Decision-1 is the short name for Microsoft-Decision-1, a model Microsoft released on 9 October 2026 that never writes a sentence. You give it a situation, a question and a fixed list of answers, and it returns a calibrated probability for each answer in a single pass. Satya Nadella introduced it on X as “our new model for fast decision-making”, and Microsoft prices it at $0.042 per million input tokens, with output tokens free.
The same day, Google began rolling out thinking levels in the Gemini app. A new Low, Medium and High effort setting replaces the old “Extended” option, and free users lost the model picker altogether, moving to an “Auto” mode that sends most prompts to Gemini Flash-Lite.
The two launches look unrelated, but they answer the same question: how much computing power should a task get? Decision-1 answers it by swapping a general chatbot for a small specialist model when the answer is a closed choice. Gemini thinking levels answer it by letting you turn the reasoning effort of one general model up or down. This article explains both, checks Microsoft’s benchmark claims against rival developers’ objections and against an earlier version of Microsoft’s own blog post, and works through what each option costs.
Table of contents
- What Microsoft Launched With Decision-1
- Decision-1 Benchmarks: What Microsoft Measured
- The Quiet Edit to Decision-1’s Speed Claim
- What Google Changed With Gemini Thinking Levels
- Gemini Thinking Levels in the API and Your Bill
- Decision-1 vs Gemini Thinking Levels: Two Ways to Spend Less Compute
- What Decision-1 Costs per Million Decisions
- What Decision-1 and Gemini Thinking Levels Mean for UK Businesses
- Decision-1 and Gemini Thinking Levels FAQ
- References and Further Reading
What Microsoft Launched With Decision-1
Microsoft calls decision models “an important new category in AI”. In its launch announcement, it says that unlike a large language model, which is “designed to generate text or reason through complex problems”, a decision model is “purpose-built to deliver structured outputs that software can immediately act on”.
What a decision model returns
The model handles yes/no questions, multiple-choice questions and ratings. It can also grade an AI response, or a proposed agent action, against a rubric. In every case the output is a probability per option, not prose, and Microsoft treats that number as part of the product. “A 90% prediction should be right about nine times out of 10 on representative cases,” the blog says, so an application can use the score to decide when to act, when to defer and when to ask a person to review.
That makes it a component rather than an assistant. Microsoft’s list of suggested uses includes model routing, incident response routing, intent analysis, data labelling, AI judging, search relevance, content filtering, code scanning, robotics and screening hypotheses in scientific work. For agent builders, the headline use is “agent controls”: checking an agent’s proposed next step and deciding whether it should continue, stop, retry or hand off to a model, a tool or a human.
How Microsoft built the model
Microsoft post-trained Alibaba’s open-weight Qwen3.5-9B for single-pass scoring, and says it “will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI”. The OpenRouter model card adds that the weights are “updated continually while the API shape stays the same”, and that the model is not intended for open-ended generation, conversation, translation or summarisation.
Unlike several rival decision models, the weights are not public. It is a hosted API with a 32,768-token context window. It is listed as a “Direct from Azure” model in the Microsoft Foundry catalogue and is also served through OpenRouter. There it runs on a separate Decisions API, so ordinary chat-completions SDKs will not work with it.
What Microsoft is already doing with it
Nadella’s launch post said Microsoft is “already testing it across Microsoft for everything from incident response and quality control to scientific discovery”. The blog gives four internal examples:
- Xbox Research sorted more than 10,000 pieces of player feedback from surveys, Steam and X into themes its researchers had defined. The model was “competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive”.
- The Copilot team, which scores the quality of chat and agent responses, found it competitive with GPT-5.6 Luna and 100 times faster.
- On-call engineers used it to retrieve knowledge during live incidents, where Microsoft says it “performed better and faster than an LLM”.
- Microsoft Discovery used it to grade experiments in an adaptive replanning loop. Its scores were 46 times more consistent than an LLM’s, at three times the speed, and replanning ran nearly four times faster.
These are Microsoft’s figures from Microsoft’s own teams. None has been published with a method or a dataset, so treat these case studies as claims to test, not results to plan around.
Decision-1 Benchmarks: What Microsoft Measured
Microsoft says Decision-1 achieved “the highest accuracy in our 36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training”. MarkTechPost, which read the figures off Microsoft’s charts, puts the exact count at 147,137 questions across nine systems, with Microsoft’s model averaging 83.5% accuracy.
Accuracy, calibration and speed
The table sets out the four systems MarkTechPost tabulated from Microsoft’s charts. Calibration measures whether a model’s stated probabilities match how often it is right, so a higher score is better.
| Measure | Microsoft-Decision-1 | Quyet-1.0-Large | H2O-Lightning-4B | GPT-6 Luna (OpenAI Decisions) |
|---|---|---|---|---|
| Developer | Microsoft | Chinh Nguyen | H2O.ai | OpenAI |
| Base model | Qwen3.5-9B | Gemma-4-31B-it | Qwen3.5-4B | Not disclosed |
| Weights | Closed, API only | Apache-2.0 | Apache-2.0 | Closed, API only |
| Average accuracy, 36 benchmarks | 83.5% | 81.9% | 77.2% | 79.4% |
| Calibration score | 92.2 | 93.1 | 91.8 | 89.9 |
| Median latency in Microsoft’s chart | 85 ms | 380 ms | 210 ms | 300 ms |
| Price | $0.042 per million input tokens, output free | Self-hosted | Self-hosted | $0.10 per million input tokens, output free |
Two details stand out. Microsoft’s model leads on accuracy by 1.6 points over Quyet-1.0-Large, a 31-billion-parameter open model published on Hugging Face by developer Chinh Nguyen, but Quyet beats it on calibration. And the gap between the best and worst of these four systems is 6.3 points, which is smaller than the launch language suggests.
Robustness and safety tests
Microsoft also tested whether equivalent requests produce equivalent decisions. It rewrote each request eight ways, including paraphrasing it and shuffling or reversing the answer options. The model changed its answer on 1.3% of these perturbations on average, and never when only the order or wording of the options changed.
For safety, Microsoft ran 5,250 requests across 11 benchmarks covering harmful content, jailbreak attempts and prompt injection. It says the model “successfully refused harmful behavior while retaining a high degree of utility”, but it has not published a score. For a model pitched at screening risky agent actions and cybersecurity triage, that is the number buyers will most want to see.
The Quiet Edit to Decision-1's Speed Claim
The speed claim is where the launch got messy. When the blog post went live on 9 October, it said Decision-1 was “4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol”. That wording survives in an Internet Archive snapshot taken at 18:39 UTC, about two minutes after Nadella’s post, and TestingCatalog quoted it that evening.
The live page now reads differently. It says the model was “2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up”, and it carries an editor’s note: “This post was updated from the original to add benchmarks for Jev on accuracy and calibration.” The note does not mention that the named runner-up also changed.
What H2O.ai says Microsoft got wrong
The likely trigger is on H2O.ai’s model card for H2O-Lightning-4B, which now carries a footnote aimed at Microsoft’s chart. H2O says that for self-hosted models, Microsoft used the JevBench leaderboard’s “adjusted” latency, which is the measured time multiplied by two plus 0.15 seconds. The board itself labels that formula “assumption, not measured”. On that basis, H2O says, its model “appeared at 210 ms instead of its measured 29 ms”.
The arithmetic checks out: 29 ms doubled is 58 ms, and adding 150 ms gives 208 ms, which rounds to the 210 ms in Microsoft’s chart. H2O’s corrected chart also puts Quyet-1.0-Large at 116 ms measured, against the 380 ms Microsoft showed. Divide Quyet’s 380 ms by Microsoft’s 85 ms and you recover the original claim: 4.47, the “4.5 times”. Against H2O’s adjusted 210 ms, the multiple is 2.47, the “2.5 times” on the live page.
Bar widths are each figure divided by 380 ms. GPT-6 Sol, at 3,010 ms in Microsoft’s chart, is off this scale.
Why the 35x claim needs context
The 35-times figure compares 85 ms with GPT-6 Sol’s 3.01 seconds: 3,010 divided by 85 is 35.4. That pits a model which reads its input once and emits a probability against a frontier chat model that generates text. It shows the gap between the two kinds of model rather than anything unique to Microsoft’s model, because the other dedicated decision models in the same chart sit between 210 and 380 ms.
The figures are also measured differently. Microsoft timed its own model through Foundry, while H2O’s 29 ms is a raw timing on one H100 GPU. OpenRouter, which forwards every Decision-1 request to Azure, shows a median of 0.27 seconds end to end, network included. None of these numbers is wrong; they measure different things. Microsoft’s model is also absent from the JevBench leaderboard itself, where H2O-Lightning-4B v1.1 led the open-weight composite on 7 October with 72.5, ahead of Jev at 71.5.
Why the latency argument matters
Microsoft makes the case for speed itself: “adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow.” If your agent chains decisions, a vendor’s latency figure turns directly into how long your users wait. That is why the method behind each number matters more than the multiple on the slide.
What Google Changed With Gemini Thinking Levels
Google’s change reached the Gemini app on the evening of 9 October, Pacific time, and was first reported in detail by 9to5Google. It has two parts: new effort settings for everyone, and a new default for people who do not pay.
Low, Medium and High replace Extended
Until this week, the Gemini app offered one extra reasoning switch, “Extended thinking”, described as being for complex problem solving. Three effort levels replace it, shown in the app as:
- Low: “Quick and efficient”
- Medium: “Balanced depth”
- High: “Extra thorough”
Google’s help page on Gemini model access says that “for each available model, you can also select the model’s effort level from low, medium, and high”. It states the trade-off plainly: a higher level “will increase the model’s ability to complete tasks and provide more thorough answers to queries, but will also use more of your limit”. According to 9to5Google, the three names match the thinking levels in Google AI Studio and the Antigravity coding tool, and the rollout to subscribers is still under way.
Free users move to Auto
The second part is the one users notice first. People “without a plan” can no longer choose a model. Their only option is Auto, which Google explains this way: “Most prompts will go to Flash-Lite, our fastest model. Prompts that need deeper reasoning may go to Flash or Pro.” If a free user turns off smart model selection, every response comes from Flash-Lite.
Google announced this in advance, and we covered the warning in our article on Google’s changes to Gemini model access. What is new is that it is now live on most free accounts, and that it arrives alongside the effort control. In effect, Google is doing for free users what Decision-1 is designed to let developers do: route each request to the cheapest model that can handle it.
| Plan | US price per month | Models after the October changes | Usage allowance |
|---|---|---|---|
| Without a plan | Free | Auto: mostly Flash-Lite, sometimes Flash or Pro | Standard limits |
| Google AI Plus | $4.99 | Flash-Lite and Flash; Pro being removed | 2x standard |
| Google AI Pro | $19.99 | Flash-Lite, Flash and Pro; Deep Think being added | 4x standard |
| Google AI Ultra | $99.99 or $199.99 | Flash-Lite, Flash and Pro, plus Deep Think | 5x or 20x AI Pro |
Sources: Google’s help page, Google’s US subscriptions page and 9to5Google. AI Plus subscribers will receive an email before their change takes effect.
Why effort now costs you allowance
Since 17 May 2026, the Gemini app has used compute-based usage limits that refresh every five hours until you reach a weekly cap. Google’s help page says the calculation factors in “the complexity of your prompt, the features you use, and the length of your chat”, and it already lists “Extended thinking and Deep Think” among the premium features that use up allowance faster. Gemini thinking levels make that trade-off explicit: High buys a more careful answer with more of your five-hour budget.
Gemini Thinking Levels in the API and Your Bill
Developers have had this control for longer. The Gemini API’s thinking guide, last updated on 9 October, exposes a thinking_level parameter with four values: minimal, low, medium and high. Each model has its own default and its own supported range.
| Model | Default thinking | Levels supported | Output price per million tokens, thinking included |
|---|---|---|---|
| gemini-3.5-flash-lite | Minimal | Minimal, low, medium, high | $2.50 |
| gemini-3.8-flash | Medium | Low, medium, high | $3.75 until 31 December 2026 |
| gemini-3.6-flash | Medium | Minimal, low, medium, high | $3.75 until 31 December 2026 |
| gemini-3.1-pro-preview | High | Low, medium, high | $12.00 for prompts up to 200k tokens |
Sources: the Gemini API thinking guide and pricing page, both last updated on 9 October 2026. Both Flash prices double from 1 January 2027.
Thinking tokens are billed as output
Google’s pricing page lists output prices as “including thinking tokens”. Every token a model spends reasoning before it answers is billed at the output rate, which is the expensive one. On Gemini 3.8 Flash, the $3.75 output rate is five times the $0.75 input rate. Decision-1 sits at the other extreme, with output priced at zero.
The defaults matter. A developer who calls Gemini 3.8 Flash without setting a level gets medium thinking, and pays for it; 3.1 Pro defaults to high. Google’s guidance names the right lever: to cut cost or latency, “lower thinking_level (low or medium) instead of setting a small max_output_tokens”. A hard token cap can stop the model mid-thought and return truncated or empty output, while still billing for the thinking it did.
What 100 thinking tokens cost at scale
Because thinking is billed per token, small habits multiply. Every 100 thinking tokens per call costs $250 per million calls on Flash-Lite, $375 on 3.8 Flash and $1,200 on 3.1 Pro: 100 million tokens multiplied by each model’s output price. A routine task left at high when low would do can cost more in thinking than in everything else.
Decision-1 vs Gemini Thinking Levels: Two Ways to Spend Less Compute
Put side by side, the two launches are different tools for one job. Both stop you paying frontier-model prices for work that does not need a frontier model, but they work at different layers.
| Question | Microsoft-Decision-1 | Gemini thinking levels |
|---|---|---|
| What is it? | A specialist decision model | An effort setting on general models |
| Who sets it? | The developer, per call | The user in the app, or the developer per call |
| What comes back? | A probability for each fixed option | Free text, code or images |
| How does it save money? | Smaller model, input-only pricing | Fewer thinking tokens, or less allowance used |
| Best for | Routing, labelling, approvals, guardrails | Drafting, analysis, coding, research |
| Main risk | Missing options or a badly framed question | Under-thinking a hard task, or paying for too much thinking on an easy one |
| Where to get it | Microsoft Foundry, OpenRouter | Gemini app, AI Studio, Gemini API, Antigravity |
When the answer is a closed set
If the right output is one of a known list, such as “refund, escalate or close”, a decision model is the cheaper, faster and more auditable choice. Decision-1 returns a score you can log and threshold, and it cannot invent a fourth option. Our earlier piece on the decision model layer in AI agent stacks explains why vendors from Cloudflare to OpenAI are carving this job out of the chatbot.
When the answer must be written
If the output is prose, code or a plan, you still need a generative model, and Gemini thinking levels are the dial that matters. Start at Low for summaries, extraction and simple replies. Move to Medium for multi-step reasoning, and keep High for genuinely hard problems where a wrong answer costs more than the extra tokens. Decision-1 has no role in that work.
Using both together
The two approaches stack. Microsoft lists model routing as one use for its model: score each incoming request, then send it to the cheapest model and effort level that can handle it. That is the same idea as Gemini’s Auto mode, which moves free users between Flash-Lite, Flash and Pro. The difference is who controls the router. In the Gemini app, Google does. In your own stack, a model such as Decision-1 lets you.
What Decision-1 Costs per Million Decisions
The most striking number in the launch is the price. At $0.042 per million input tokens, Decision-1 costs exactly what TypeSafe charges for Jev 1.13 ($42 per billion tokens, on TypeSafe’s model page). Jev’s September launch set off the current wave of copycats, which we covered in Jev’s rivals and the talk of LLM alternatives. OpenAI’s Decisions API, still in public beta, charges $0.10 per million input tokens for gpt-6-luna. Microsoft matched the incumbent’s price rather than undercutting it.
A worked example: one million support tickets
To compare like with like, assume one million classification calls, each with 600 input tokens covering the ticket, the question and the list of options. For chat models, assume a 10-token answer and no thinking. The chart is our arithmetic on each vendor’s published list price.
The working: 600 million input tokens cost $25.20 at $0.042 per million and $60.00 at $0.10. GPT-6 Luna by chat adds 10 million output tokens at $0.50 per million (its OpenRouter list price), giving $65.00. Flash-Lite costs $180 for input at $0.30 plus $25 for output at $2.50, a total of $205. Gemini 3.8 Flash costs $450 plus $37.50, or $487.50, and that assumes thinking is turned down, because its default is medium.
What the example leaves out
The chart compares price, not value. A decision model that picks the wrong option costs more than its token bill, and a frontier model may be worth its price on hard calls. Hosting is excluded too: Quyet-1.0-Large and H2O-Lightning-4B are free to download but need your own GPUs. Use the chart for the order of magnitude, then run a sample of your real decisions through each candidate before you commit. Our cost optimisation team can model this for your own volumes.
What Decision-1 and Gemini Thinking Levels Mean for UK Businesses
For most UK organisations, neither launch demands immediate action. Both, though, change how you should budget for and govern AI over the next year.
A pilot plan for Decision-1
A sensible trial takes a few weeks:
- Pick one closed decision you already make at volume, such as ticket triage or invoice approval routing, and pull 500 historical cases with known outcomes.
- Run them through Decision-1, an open model such as H2O-Lightning-4B and your current LLM. Compare accuracy, and check whether a 90% score really is right nine times in ten on your data.
- Measure latency from your own network, not from a vendor’s chart.
- Set two thresholds: act automatically above one score, and send the case to a person below the other.
Governance: automated decisions still need a human route
If a decision model’s output affects individuals, the law follows the decision, not the model. Under the Data (Use and Access) Act 2025, significant decisions based solely on automated processing still need safeguards, including the right to contest the decision and to obtain human intervention. Our guide to automated decision-making under the DUAA explains what that means in practice. A calibrated probability helps here, because it gives you a clean, logged threshold for when a person must step in.
Microsoft has not published data-residency details for Decision-1. Check where Foundry processes your requests before sending it personal data, as you would for any Azure model.
Setting Gemini thinking levels for staff
If your teams use the Gemini app, decide which level is the default for routine work. Low is enough for drafting emails and summarising documents, and it stretches the five-hour allowance further. Encourage High only for analysis where accuracy matters more than speed.
Staff on free personal accounts now get Auto with mostly Flash-Lite, so check whether anyone relies on Gemini for work without a business plan. If they do, you have a data-governance question as well as a quality one. Our AI strategy service can help set those policies.
Watch the effort trend across vendors
Effort controls are now standard. OpenAI and Anthropic both expose reasoning-effort settings in their APIs, and we covered Perplexity’s effort selector in August. Expect licence talks to shift from “which model” to “how much thinking”, and ask each vendor how effort affects your quota or bill. Decision-1 points to the other half of the same conversation: some tasks need no thinking at all.
Decision-1 and Gemini Thinking Levels FAQ
What is Microsoft Decision-1?
Decision-1, officially Microsoft-Decision-1, is a decision-scoring model released on 9 October 2026. It takes a question and a fixed set of answers and returns a calibrated probability for each, instead of writing text. It is built on Qwen3.5-9B and is available through Microsoft Foundry and OpenRouter.
How much does Decision-1 cost?
Input tokens cost $0.042 per million, and output tokens are free. That is the same price as TypeSafe’s Jev 1.13, and less than half the $0.10 per million OpenAI charges on its Decisions API.
Is Decision-1 really 35 times faster than GPT-6 Sol?
In Microsoft’s chart, its median latency is 85 ms against 3.01 seconds for GPT-6 Sol. That compares a single-pass scorer with a text-generating chat model. H2O.ai disputes how rival decision models’ latency was shown, and Microsoft changed its runner-up claim after launch.
What are Gemini thinking levels?
They are effort settings in the Gemini app: Low (“Quick and efficient”), Medium (“Balanced depth”) and High (“Extra thorough”). They replace the old Extended thinking option. Higher levels give more thorough answers but use more of your usage limit.
Can free Gemini users still choose a model?
No. Since 9 October, users without a Google AI plan are on Auto, which sends most prompts to Flash-Lite and some to Flash or Pro. With smart model selection turned off, every reply comes from Flash-Lite.
Should I use Decision-1 or Gemini thinking levels?
Use a decision model when the answer is one of a fixed list. Use Gemini thinking levels when you need written output and want to trade depth for speed or cost. Many systems will use both: a decision model to route each request, and a generative model at the right effort to answer it.
References and Further Reading
Introducing Microsoft-Decision-1, our model for fast decision-making (Microsoft)
Original 9 October version of Microsoft’s launch post (Internet Archive)
Model card in the Microsoft Foundry catalogue
Microsoft-Decision-1 (OpenRouter)
Satya Nadella’s launch post (X)
Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model (MarkTechPost)
Microsoft launches Decision-1 model in Foundry (TestingCatalog)
H2O-Lightning-4B model card (Hugging Face)
Quyet-1.0-Large model card (Hugging Face)
Jev models and pricing (TypeSafe)
Free Gemini users now on ‘Auto’ models as app adds thinking levels (9to5Google)
Gemini limiting what models free and AI Plus users can access (9to5Google)
Changes to Gemini model access and limits (Gemini Apps Help)
Gemini thinking (Gemini API documentation)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.