Decision-1 is the short name for Microsoft-Decision-1, a model Microsoft released on 9 October 2026 that never writes a sentence. You give it a situation, a question and a fixed list of answers, and it returns a calibrated probability for each answer in a single pass. Satya Nadella introduced it on X as “our new model for fast decision-making”, and Microsoft prices it at $0.042 per million input tokens, with output tokens free.

The same day, Google began rolling out thinking levels in the Gemini app. A new Low, Medium and High effort setting replaces the old “Extended” option, and free users lost the model picker altogether, moving to an “Auto” mode that sends most prompts to Gemini Flash-Lite.

The two launches look unrelated, but they answer the same question: how much computing power should a task get? Decision-1 answers it by swapping a general chatbot for a small specialist model when the answer is a closed choice. Gemini thinking levels answer it by letting you turn the reasoning effort of one general model up or down. This article explains both, checks Microsoft’s benchmark claims against rival developers’ objections and against an earlier version of Microsoft’s own blog post, and works through what each option costs.

What Microsoft Launched With Decision-1

Decision-1 - microsoft decision 1 gemini thinking levels b weather balloon on a tether above a winch

Microsoft calls decision models “an important new category in AI”. In its launch announcement, it says that unlike a large language model, which is “designed to generate text or reason through complex problems”, a decision model is “purpose-built to deliver structured outputs that software can immediately act on”.

What a decision model returns

The model handles yes/no questions, multiple-choice questions and ratings. It can also grade an AI response, or a proposed agent action, against a rubric. In every case the output is a probability per option, not prose, and Microsoft treats that number as part of the product. “A 90% prediction should be right about nine times out of 10 on representative cases,” the blog says, so an application can use the score to decide when to act, when to defer and when to ask a person to review.

That makes it a component rather than an assistant. Microsoft’s list of suggested uses includes model routing, incident response routing, intent analysis, data labelling, AI judging, search relevance, content filtering, code scanning, robotics and screening hypotheses in scientific work. For agent builders, the headline use is “agent controls”: checking an agent’s proposed next step and deciding whether it should continue, stop, retry or hand off to a model, a tool or a human.

How Microsoft built the model

Microsoft post-trained Alibaba’s open-weight Qwen3.5-9B for single-pass scoring, and says it “will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI”. The OpenRouter model card adds that the weights are “updated continually while the API shape stays the same”, and that the model is not intended for open-ended generation, conversation, translation or summarisation.

Unlike several rival decision models, the weights are not public. It is a hosted API with a 32,768-token context window. It is listed as a “Direct from Azure” model in the Microsoft Foundry catalogue and is also served through OpenRouter. There it runs on a separate Decisions API, so ordinary chat-completions SDKs will not work with it.

What Microsoft is already doing with it

Nadella’s launch post said Microsoft is “already testing it across Microsoft for everything from incident response and quality control to scientific discovery”. The blog gives four internal examples:

  • Xbox Research sorted more than 10,000 pieces of player feedback from surveys, Steam and X into themes its researchers had defined. The model was “competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive”.
  • The Copilot team, which scores the quality of chat and agent responses, found it competitive with GPT-5.6 Luna and 100 times faster.
  • On-call engineers used it to retrieve knowledge during live incidents, where Microsoft says it “performed better and faster than an LLM”.
  • Microsoft Discovery used it to grade experiments in an adaptive replanning loop. Its scores were 46 times more consistent than an LLM’s, at three times the speed, and replanning ran nearly four times faster.

These are Microsoft’s figures from Microsoft’s own teams. None has been published with a method or a dataset, so treat these case studies as claims to test, not results to plan around.

Decision-1 Benchmarks: What Microsoft Measured

microsoft decision 1 gemini thinking levels c trundle measuring wheel with a long handle

Microsoft says Decision-1 achieved “the highest accuracy in our 36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training”. MarkTechPost, which read the figures off Microsoft’s charts, puts the exact count at 147,137 questions across nine systems, with Microsoft’s model averaging 83.5% accuracy.

Accuracy, calibration and speed

The table sets out the four systems MarkTechPost tabulated from Microsoft’s charts. Calibration measures whether a model’s stated probabilities match how often it is right, so a higher score is better.

MeasureMicrosoft-Decision-1Quyet-1.0-LargeH2O-Lightning-4BGPT-6 Luna (OpenAI Decisions)
DeveloperMicrosoftChinh NguyenH2O.aiOpenAI
Base modelQwen3.5-9BGemma-4-31B-itQwen3.5-4BNot disclosed
WeightsClosed, API onlyApache-2.0Apache-2.0Closed, API only
Average accuracy, 36 benchmarks83.5%81.9%77.2%79.4%
Calibration score92.293.191.889.9
Median latency in Microsoft’s chart85 ms380 ms210 ms300 ms
Price$0.042 per million input tokens, output freeSelf-hostedSelf-hosted$0.10 per million input tokens, output free

Two details stand out. Microsoft’s model leads on accuracy by 1.6 points over Quyet-1.0-Large, a 31-billion-parameter open model published on Hugging Face by developer Chinh Nguyen, but Quyet beats it on calibration. And the gap between the best and worst of these four systems is 6.3 points, which is smaller than the launch language suggests.

Robustness and safety tests

Microsoft also tested whether equivalent requests produce equivalent decisions. It rewrote each request eight ways, including paraphrasing it and shuffling or reversing the answer options. The model changed its answer on 1.3% of these perturbations on average, and never when only the order or wording of the options changed.

For safety, Microsoft ran 5,250 requests across 11 benchmarks covering harmful content, jailbreak attempts and prompt injection. It says the model “successfully refused harmful behavior while retaining a high degree of utility”, but it has not published a score. For a model pitched at screening risky agent actions and cybersecurity triage, that is the number buyers will most want to see.

The Quiet Edit to Decision-1's Speed Claim

microsoft decision 1 gemini thinking levels d correction tape dispenser over a dark line

The speed claim is where the launch got messy. When the blog post went live on 9 October, it said Decision-1 was “4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol”. That wording survives in an Internet Archive snapshot taken at 18:39 UTC, about two minutes after Nadella’s post, and TestingCatalog quoted it that evening.

The live page now reads differently. It says the model was “2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up”, and it carries an editor’s note: “This post was updated from the original to add benchmarks for Jev on accuracy and calibration.” The note does not mention that the named runner-up also changed.

What H2O.ai says Microsoft got wrong

The likely trigger is on H2O.ai’s model card for H2O-Lightning-4B, which now carries a footnote aimed at Microsoft’s chart. H2O says that for self-hosted models, Microsoft used the JevBench leaderboard’s “adjusted” latency, which is the measured time multiplied by two plus 0.15 seconds. The board itself labels that formula “assumption, not measured”. On that basis, H2O says, its model “appeared at 210 ms instead of its measured 29 ms”.

The arithmetic checks out: 29 ms doubled is 58 ms, and adding 150 ms gives 208 ms, which rounds to the 210 ms in Microsoft’s chart. H2O’s corrected chart also puts Quyet-1.0-Large at 116 ms measured, against the 380 ms Microsoft showed. Divide Quyet’s 380 ms by Microsoft’s 85 ms and you recover the original claim: 4.47, the “4.5 times”. Against H2O’s adjusted 210 ms, the multiple is 2.47, the “2.5 times” on the live page.

Median latency of decision models, by who measured it and how (milliseconds)
Quyet-1.0-Large, Microsoft’s chart (JevBench adjusted) 380 ms
GPT-6 Luna Decisions, Microsoft’s chart 300 ms
Microsoft-Decision-1, OpenRouter end to end (3-day median) 270 ms
H2O-Lightning-4B, Microsoft’s chart (JevBench adjusted) 210 ms
Quyet-1.0-Large, measured on one H100 (H2O’s corrected chart) 116 ms
Microsoft-Decision-1, Microsoft’s chart (via Foundry) 85 ms
H2O-Lightning-4B, measured on one H100 (H2O’s model card) 29 ms

Bar widths are each figure divided by 380 ms. GPT-6 Sol, at 3,010 ms in Microsoft’s chart, is off this scale.

Why the 35x claim needs context

The 35-times figure compares 85 ms with GPT-6 Sol’s 3.01 seconds: 3,010 divided by 85 is 35.4. That pits a model which reads its input once and emits a probability against a frontier chat model that generates text. It shows the gap between the two kinds of model rather than anything unique to Microsoft’s model, because the other dedicated decision models in the same chart sit between 210 and 380 ms.

The figures are also measured differently. Microsoft timed its own model through Foundry, while H2O’s 29 ms is a raw timing on one H100 GPU. OpenRouter, which forwards every Decision-1 request to Azure, shows a median of 0.27 seconds end to end, network included. None of these numbers is wrong; they measure different things. Microsoft’s model is also absent from the JevBench leaderboard itself, where H2O-Lightning-4B v1.1 led the open-weight composite on 7 October with 72.5, ahead of Jev at 71.5.

Why the latency argument matters

Microsoft makes the case for speed itself: “adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow.” If your agent chains decisions, a vendor’s latency figure turns directly into how long your users wait. That is why the method behind each number matters more than the multiple on the slide.

What Google Changed With Gemini Thinking Levels

microsoft decision 1 gemini thinking levels e three measuring cups in a stepped row

Google’s change reached the Gemini app on the evening of 9 October, Pacific time, and was first reported in detail by 9to5Google. It has two parts: new effort settings for everyone, and a new default for people who do not pay.

Low, Medium and High replace Extended

Until this week, the Gemini app offered one extra reasoning switch, “Extended thinking”, described as being for complex problem solving. Three effort levels replace it, shown in the app as:

  • Low: “Quick and efficient”
  • Medium: “Balanced depth”
  • High: “Extra thorough”

Google’s help page on Gemini model access says that “for each available model, you can also select the model’s effort level from low, medium, and high”. It states the trade-off plainly: a higher level “will increase the model’s ability to complete tasks and provide more thorough answers to queries, but will also use more of your limit”. According to 9to5Google, the three names match the thinking levels in Google AI Studio and the Antigravity coding tool, and the rollout to subscribers is still under way.

Free users move to Auto

The second part is the one users notice first. People “without a plan” can no longer choose a model. Their only option is Auto, which Google explains this way: “Most prompts will go to Flash-Lite, our fastest model. Prompts that need deeper reasoning may go to Flash or Pro.” If a free user turns off smart model selection, every response comes from Flash-Lite.

Google announced this in advance, and we covered the warning in our article on Google’s changes to Gemini model access. What is new is that it is now live on most free accounts, and that it arrives alongside the effort control. In effect, Google is doing for free users what Decision-1 is designed to let developers do: route each request to the cheapest model that can handle it.

PlanUS price per monthModels after the October changesUsage allowance
Without a planFreeAuto: mostly Flash-Lite, sometimes Flash or ProStandard limits
Google AI Plus$4.99Flash-Lite and Flash; Pro being removed2x standard
Google AI Pro$19.99Flash-Lite, Flash and Pro; Deep Think being added4x standard
Google AI Ultra$99.99 or $199.99Flash-Lite, Flash and Pro, plus Deep Think5x or 20x AI Pro

Sources: Google’s help page, Google’s US subscriptions page and 9to5Google. AI Plus subscribers will receive an email before their change takes effect.

Why effort now costs you allowance

Since 17 May 2026, the Gemini app has used compute-based usage limits that refresh every five hours until you reach a weekly cap. Google’s help page says the calculation factors in “the complexity of your prompt, the features you use, and the length of your chat”, and it already lists “Extended thinking and Deep Think” among the premium features that use up allowance faster. Gemini thinking levels make that trade-off explicit: High buys a more careful answer with more of your five-hour budget.

Gemini Thinking Levels in the API and Your Bill

microsoft decision 1 gemini thinking levels f sugar cube stack on a saucer

Developers have had this control for longer. The Gemini API’s thinking guide, last updated on 9 October, exposes a thinking_level parameter with four values: minimal, low, medium and high. Each model has its own default and its own supported range.

ModelDefault thinkingLevels supportedOutput price per million tokens, thinking included
gemini-3.5-flash-liteMinimalMinimal, low, medium, high$2.50
gemini-3.8-flashMediumLow, medium, high$3.75 until 31 December 2026
gemini-3.6-flashMediumMinimal, low, medium, high$3.75 until 31 December 2026
gemini-3.1-pro-previewHighLow, medium, high$12.00 for prompts up to 200k tokens

Sources: the Gemini API thinking guide and pricing page, both last updated on 9 October 2026. Both Flash prices double from 1 January 2027.

Thinking tokens are billed as output

Google’s pricing page lists output prices as “including thinking tokens”. Every token a model spends reasoning before it answers is billed at the output rate, which is the expensive one. On Gemini 3.8 Flash, the $3.75 output rate is five times the $0.75 input rate. Decision-1 sits at the other extreme, with output priced at zero.

The defaults matter. A developer who calls Gemini 3.8 Flash without setting a level gets medium thinking, and pays for it; 3.1 Pro defaults to high. Google’s guidance names the right lever: to cut cost or latency, “lower thinking_level (low or medium) instead of setting a small max_output_tokens”. A hard token cap can stop the model mid-thought and return truncated or empty output, while still billing for the thinking it did.

What 100 thinking tokens cost at scale

Because thinking is billed per token, small habits multiply. Every 100 thinking tokens per call costs $250 per million calls on Flash-Lite, $375 on 3.8 Flash and $1,200 on 3.1 Pro: 100 million tokens multiplied by each model’s output price. A routine task left at high when low would do can cost more in thinking than in everything else.

Decision-1 vs Gemini Thinking Levels: Two Ways to Spend Less Compute

Put side by side, the two launches are different tools for one job. Both stop you paying frontier-model prices for work that does not need a frontier model, but they work at different layers.

QuestionMicrosoft-Decision-1Gemini thinking levels
What is it?A specialist decision modelAn effort setting on general models
Who sets it?The developer, per callThe user in the app, or the developer per call
What comes back?A probability for each fixed optionFree text, code or images
How does it save money?Smaller model, input-only pricingFewer thinking tokens, or less allowance used
Best forRouting, labelling, approvals, guardrailsDrafting, analysis, coding, research
Main riskMissing options or a badly framed questionUnder-thinking a hard task, or paying for too much thinking on an easy one
Where to get itMicrosoft Foundry, OpenRouterGemini app, AI Studio, Gemini API, Antigravity

When the answer is a closed set

If the right output is one of a known list, such as “refund, escalate or close”, a decision model is the cheaper, faster and more auditable choice. Decision-1 returns a score you can log and threshold, and it cannot invent a fourth option. Our earlier piece on the decision model layer in AI agent stacks explains why vendors from Cloudflare to OpenAI are carving this job out of the chatbot.

When the answer must be written

If the output is prose, code or a plan, you still need a generative model, and Gemini thinking levels are the dial that matters. Start at Low for summaries, extraction and simple replies. Move to Medium for multi-step reasoning, and keep High for genuinely hard problems where a wrong answer costs more than the extra tokens. Decision-1 has no role in that work.

Using both together

The two approaches stack. Microsoft lists model routing as one use for its model: score each incoming request, then send it to the cheapest model and effort level that can handle it. That is the same idea as Gemini’s Auto mode, which moves free users between Flash-Lite, Flash and Pro. The difference is who controls the router. In the Gemini app, Google does. In your own stack, a model such as Decision-1 lets you.

What Decision-1 Costs per Million Decisions

The most striking number in the launch is the price. At $0.042 per million input tokens, Decision-1 costs exactly what TypeSafe charges for Jev 1.13 ($42 per billion tokens, on TypeSafe’s model page). Jev’s September launch set off the current wave of copycats, which we covered in Jev’s rivals and the talk of LLM alternatives. OpenAI’s Decisions API, still in public beta, charges $0.10 per million input tokens for gpt-6-luna. Microsoft matched the incumbent’s price rather than undercutting it.

A worked example: one million support tickets

To compare like with like, assume one million classification calls, each with 600 input tokens covering the ticket, the question and the list of options. For chat models, assume a 10-token answer and no thinking. The chart is our arithmetic on each vendor’s published list price.

Cost of one million 600-token classification calls, US dollars (list prices, our arithmetic)
Gemini 3.8 Flash, thinking turned down $487.50
Gemini 3.5 Flash-Lite, minimal thinking $205.00
GPT-6 Luna via chat, 10-token answer $65.00
OpenAI Decisions API, gpt-6-luna $60.00
TypeSafe Jev 1.13 $25.20
Microsoft-Decision-1 $25.20

The working: 600 million input tokens cost $25.20 at $0.042 per million and $60.00 at $0.10. GPT-6 Luna by chat adds 10 million output tokens at $0.50 per million (its OpenRouter list price), giving $65.00. Flash-Lite costs $180 for input at $0.30 plus $25 for output at $2.50, a total of $205. Gemini 3.8 Flash costs $450 plus $37.50, or $487.50, and that assumes thinking is turned down, because its default is medium.

What the example leaves out

The chart compares price, not value. A decision model that picks the wrong option costs more than its token bill, and a frontier model may be worth its price on hard calls. Hosting is excluded too: Quyet-1.0-Large and H2O-Lightning-4B are free to download but need your own GPUs. Use the chart for the order of magnitude, then run a sample of your real decisions through each candidate before you commit. Our cost optimisation team can model this for your own volumes.

What Decision-1 and Gemini Thinking Levels Mean for UK Businesses

For most UK organisations, neither launch demands immediate action. Both, though, change how you should budget for and govern AI over the next year.

A pilot plan for Decision-1

A sensible trial takes a few weeks:

  1. Pick one closed decision you already make at volume, such as ticket triage or invoice approval routing, and pull 500 historical cases with known outcomes.
  2. Run them through Decision-1, an open model such as H2O-Lightning-4B and your current LLM. Compare accuracy, and check whether a 90% score really is right nine times in ten on your data.
  3. Measure latency from your own network, not from a vendor’s chart.
  4. Set two thresholds: act automatically above one score, and send the case to a person below the other.

Governance: automated decisions still need a human route

If a decision model’s output affects individuals, the law follows the decision, not the model. Under the Data (Use and Access) Act 2025, significant decisions based solely on automated processing still need safeguards, including the right to contest the decision and to obtain human intervention. Our guide to automated decision-making under the DUAA explains what that means in practice. A calibrated probability helps here, because it gives you a clean, logged threshold for when a person must step in.

Microsoft has not published data-residency details for Decision-1. Check where Foundry processes your requests before sending it personal data, as you would for any Azure model.

Setting Gemini thinking levels for staff

If your teams use the Gemini app, decide which level is the default for routine work. Low is enough for drafting emails and summarising documents, and it stretches the five-hour allowance further. Encourage High only for analysis where accuracy matters more than speed.

Staff on free personal accounts now get Auto with mostly Flash-Lite, so check whether anyone relies on Gemini for work without a business plan. If they do, you have a data-governance question as well as a quality one. Our AI strategy service can help set those policies.

Watch the effort trend across vendors

Effort controls are now standard. OpenAI and Anthropic both expose reasoning-effort settings in their APIs, and we covered Perplexity’s effort selector in August. Expect licence talks to shift from “which model” to “how much thinking”, and ask each vendor how effort affects your quota or bill. Decision-1 points to the other half of the same conversation: some tasks need no thinking at all.

Decision-1 and Gemini Thinking Levels FAQ

What is Microsoft Decision-1?

Decision-1, officially Microsoft-Decision-1, is a decision-scoring model released on 9 October 2026. It takes a question and a fixed set of answers and returns a calibrated probability for each, instead of writing text. It is built on Qwen3.5-9B and is available through Microsoft Foundry and OpenRouter.

How much does Decision-1 cost?

Input tokens cost $0.042 per million, and output tokens are free. That is the same price as TypeSafe’s Jev 1.13, and less than half the $0.10 per million OpenAI charges on its Decisions API.

Is Decision-1 really 35 times faster than GPT-6 Sol?

In Microsoft’s chart, its median latency is 85 ms against 3.01 seconds for GPT-6 Sol. That compares a single-pass scorer with a text-generating chat model. H2O.ai disputes how rival decision models’ latency was shown, and Microsoft changed its runner-up claim after launch.

What are Gemini thinking levels?

They are effort settings in the Gemini app: Low (“Quick and efficient”), Medium (“Balanced depth”) and High (“Extra thorough”). They replace the old Extended thinking option. Higher levels give more thorough answers but use more of your usage limit.

Can free Gemini users still choose a model?

No. Since 9 October, users without a Google AI plan are on Auto, which sends most prompts to Flash-Lite and some to Flash or Pro. With smart model selection turned off, every reply comes from Flash-Lite.

Should I use Decision-1 or Gemini thinking levels?

Use a decision model when the answer is one of a fixed list. Use Gemini thinking levels when you need written output and want to trade depth for speed or cost. Many systems will use both: a decision model to route each request, and a generative model at the right effort to answer it.

References and Further Reading