LLM API pricing in August 2026 looks calmer on the surface than at any point in the platform era — and is more chaotic underneath than ever. Sticker rates for artificial intelligence models have converged into recognisable tiers, yet the real cost of running an identical workload can still differ by orders of magnitude depending on how, when and where you send your tokens. This guide, part of our AI models and tools hub, gathers every current rate from the seven major providers into one place.
Three fresh moves make this the right week to re-check your budget. Anthropic has made Claude Sonnet 5’s $2 input / $10 output rate permanent — a price first framed as an introductory offer in the official Claude Sonnet 5 launch announcement. DeepSeek switches to peak and off-peak billing from 16 August. And Google’s promotional rate on Gemini 3.7 Flash runs only until 31 December 2026. If you set your model budgets in spring, your assumptions are already stale.
We focus strictly on pay-per-token developer rates here, not chat subscriptions — for consumer chat apps, see our separate AI comparison. Because LLM API pricing changes so quickly, treat every table below as a snapshot: all figures were verified against official provider price lists in August 2026, with two flagged exceptions for Alibaba where third-party trackers had to stand in.
The guide opens with how LLM API pricing actually works, presents a master comparison table across OpenAI, Anthropic, Google, Mistral, DeepSeek, Qwen and xAI, then walks provider by provider before turning to the discount stack — caching, batch and off-peak — where most real savings now live.
Table of contents
- Why August 2026 Is a Watershed for LLM API Pricing
- How LLM API Pricing Works: Tokens, Caching and Tiers
- The Master LLM API Pricing Table (August 2026)
- OpenAI Rates: GPT-5.6, GPT-5.5 and Legacy Models
- Anthropic Claude LLM API Pricing
- Google Gemini Rates and the Promotional Clock
- Mistral Rates: Open-Weight Economics
- DeepSeek Rates: The Price Floor of 2026
- Alibaba Qwen Rates — and a Sourcing Caveat
- xAI Grok Rates: Simple Tokens, Separate Tools
- Stacking Discounts on LLM API Pricing
- Choosing the Best Rates for Your Workload
- A Practical Playbook to Cut Your LLM API Bill
- The Bottom Line for August 2026
- FAQ: LLM API Pricing Questions Answered
- References
Why August 2026 Is a Watershed for LLM API Pricing
The 640-fold spread in one number
Across the seven providers covered here, flagship-tier output tokens span $0.28 per million (DeepSeek V4-Flash) to $180 per million (OpenAI’s gpt-5.5-pro). That is a roughly 640-fold spread between the cheapest and dearest output token on any current price list. The same comparison notes that every provider except xAI now offers a 50% batch discount, which means the effective spread widens further once discounts are applied.
Three fresh moves reshaping LLM API pricing
Three dated events define this month’s LLM API pricing picture. First, Anthropic confirmed that Claude Sonnet 5’s $2/$10 introductory rate is now the standard price — the scheduled increase to $3/$15 on 1 September 2026 will not occur. Second, from 16:00 UTC on 16 August 2026 DeepSeek moves to peak/off-peak billing, with off-peak rates at 50% of peak. Third, Gemini 3.7 Flash’s $0.75/$3.75 promotional rate expires on 31 December 2026.
Why sticker rates mislead
Published input and output rates are only the opening bid. Caching multipliers knock 90–98% off repeated input, batch processing halves both directions, and time-of-day billing is about to halve DeepSeek again. Any LLM API pricing comparison that stops at the sticker will rank providers in the wrong order for your workload. That is why this guide treats the discount stack as a first-class topic rather than a footnote.
How LLM API Pricing Works: Tokens, Caching and Tiers
Input and output tokens in LLM API pricing
Every provider in this guide bills the same two meters: input tokens (your prompt, system message, documents and tool results) and output tokens (the model’s reply). Output is almost always dearer — anywhere from 2x input at DeepSeek V4-Flash ($0.14 in, $0.28 out) to 6x at OpenAI’s gpt-5.6-sol ($5 in, $30 out). Understanding your own input-to-output ratio is the first step in any LLM API pricing analysis, because a summarisation workload and a generation workload favour entirely different price lists.
Tokenizers quietly change your effective rate
Per-token rates are not the whole story, because providers tokenize text differently. Anthropic’s documentation states that Claude 4.7-and-later models use a newer tokenizer that produces approximately 30% more tokens for the same text than Sonnet 4.6 and earlier. Per-token prices are unchanged, but the effective cost of a given request shifts. When comparing LLM API pricing across vendors, always compare cost per request on your own traffic, not cost per token in the abstract.
Context tiers and long-context surcharges
Several price lists step up with prompt length. Gemini 3.1 Pro Preview charges $2.00/$12.00 per million up to 200k input tokens, rising to $4.00/$18.00 above that. xAI doubles Grok 4.6 from $2.00/$6.00 to $4.00/$12.00 for the whole request once a prompt reaches 200k tokens. Anthropic goes the other way: Claude 4.6 and later include the full 1M-token context window at standard rates, with no long-context premium.
The $2 convergence club
One striking feature of current LLM API pricing is how many flagships now share a sticker. Claude Sonnet 5, gpt-5.6-terra, Qwen3.8-Max and Grok 4.6 all charge exactly $2.00 per million input tokens, while their output rates fan out from $6 (Qwen, Grok) through $10 (Sonnet 5) to $12 (terra). When input converges like this, the deciding factors move to output rates, caching mechanics and quality — which is precisely how the rest of this guide compares them.
The discount stack in brief
Three levers sit under the sticker: prompt caching (a discount on input the provider has already seen), batch APIs (a 50% discount for asynchronous processing) and, newly at DeepSeek, off-peak windows. Stacked well, they routinely cut real bills by 50–90% — the deep-dive section below walks each lever provider by provider.
The Master LLM API Pricing Table (August 2026)
One table, seven providers, twenty-two models: this is current LLM API pricing per million tokens, drawn from the official price lists in the References (Qwen rows from third-party trackers, as flagged later).
| Provider | Model | Input $/M | Output $/M | Cached input $/M |
|---|---|---|---|---|
| OpenAI | gpt-5.5-pro | $30.00 | $180.00 | — |
| OpenAI | gpt-5.6-sol | $5.00 | $30.00 | $0.50 |
| OpenAI | gpt-5.6-terra | $2.00 | $12.00 | $0.20 |
| OpenAI | gpt-5.6-luna | $0.20 | $1.20 | $0.02 |
| Anthropic | Claude Fable 5 | $10.00 | $50.00 | $1.00 |
| Anthropic | Claude Opus 5 | $5.00 | $25.00 | $0.50 (0.1x) |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | $0.20 (0.1x) |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 (0.1x) |
| Gemini 3.1 Pro Preview | $2.00–$4.00 | $12.00–$18.00 | — | |
| Gemini 3.5 Flash | $1.50 | $9.00 | — | |
| Gemini 3.7 Flash (promo) | $0.75 | $3.75 | $0.075 + storage | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | — | |
| Mistral | Mistral Medium 3.5 | $1.50 | $7.50 | up to 90% off |
| Mistral | Mistral Large 3 | $0.50 | $1.50 | up to 90% off |
| Mistral | Mistral Small 4 | $0.15 | $0.60 | up to 90% off |
| DeepSeek | deepseek-v4-pro | $0.435 | $0.87 | $0.003625 |
| DeepSeek | deepseek-v4-flash | $0.14 | $0.28 | $0.0028 |
| Alibaba | Qwen3.8-Max | $2.00 | $6.00 | $0.25 |
| Alibaba | Qwen3.5-Plus | $0.40 | $2.40 | — |
| Alibaba | Qwen3.5-Flash | $0.10 | $0.40 | — |
| xAI | Grok 4.6 | $2.00 | $6.00 | $0.50 |
| xAI | Grok 4.3 | $1.25 | $2.50 | $0.20 |
How to read this LLM API pricing table
Rates are per million tokens: standard tier, sub-200k prompts, cache misses. Ranged cells (Gemini 3.1 Pro Preview) show the sub-200k and above-200k tiers. Anthropic cache reads are listed at 0.1x base input as its price list specifies. Gemini 3.7 Flash caching adds $0.50 per hour of storage on top of the $0.075 read rate. This LLM API pricing table deliberately excludes fine-tuning, embeddings and image generation — text in, text out only.
Flagship output rates, charted
The chart below plots the output rates for the five dearest flagship-tier entries in the table above — gpt-5.5-pro’s $180 per million dwarfs everything else in current LLM API pricing.
Tier winners at a glance
Reading down the table, the pattern is clear. In the premium tier, Claude Opus 5 at $5/$25 undercuts gpt-5.6-sol at $5/$30 on output. In the workhorse tier, Claude Sonnet 5 and gpt-5.6-terra sit at an identical $2 input, with Sonnet 5’s $10 output edging terra’s $12. At the floor, DeepSeek V4-Flash’s $0.14/$0.28 leads every list. Whoever your incumbent is, current LLM API pricing gives you a credible alternative one row away.
OpenAI Rates: GPT-5.6, GPT-5.5 and Legacy Models
GPT-5.6 Sol, Terra and Luna LLM API pricing
OpenAI’s current flagship family is GPT-5.6, sold in three sizes. gpt-5.6-sol costs $5.00 input / $30.00 output per million tokens, with cached input at $0.50. gpt-5.6-terra is the mid-tier at $2.00/$12.00 (cached $0.20), and gpt-5.6-luna the small tier at $0.20/$1.20 (cached $0.02). The trio spans a 25x price range within one family — evidence of how much LLM API pricing now happens inside a provider rather than between providers.
GPT-5.5, gpt-5.5-pro and the $180 output rate
GPT-5.5, released on 23 April 2026, is now marked legacy on OpenAI’s price list, superseded by the GPT-5.6 family. It remains listed at $5.00/$30.00 with cached input at $0.50 — identical to gpt-5.6-sol. The outlier is gpt-5.5-pro at $30.00 input / $180.00 output per million: the most expensive mainstream flagship output on any current price list. For what the 5.5 generation changed in practice, see our earlier look at GPT-5.5 Instant and shopping intent.
Older OpenAI tiers still listed
OpenAI keeps its older generations purchasable, and they matter for budget work. gpt-5.4 is listed at $2.50/$15.00, gpt-5.4-mini at $0.75/$4.50 and gpt-5.4-nano at $0.20/$1.25. A generation further back, gpt-5 costs $1.25/$10.00, gpt-5-mini $0.25/$2.00 and gpt-5-nano just $0.05/$0.40 per million tokens — the cheapest input rate anywhere in this guide.
OpenAI caching and batch discounts
OpenAI’s Batch API prices are 50% of standard rates: gpt-5.6-sol drops to $2.50 input / $15.00 output per million in batch. Cached input is typically 90% off the standard input rate, applied automatically when a prompt prefix repeats. Combine the two on a cache-heavy batch workload and sol’s effective input cost falls from $5.00 towards the $0.25–$0.50 region — a reminder that OpenAI LLM API pricing rewards architecture as much as model choice.
Anthropic Claude LLM API Pricing
Claude Opus 5 and the wider Opus line
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens. Notably, Anthropic prices Opus 4.8, 4.7, 4.6 and 4.5 identically — upgrading generations does not change the rate. At $25 output, Opus 5 undercuts OpenAI’s $30 gpt-5.6-sol while sitting well below Anthropic’s own Fable tier.
Claude Sonnet 5: $2/$10 is now the standard rate
Claude Sonnet 5’s $2 input / $10 output rate was announced at launch as introductory pricing through 31 August 2026. Anthropic has since made it the standard price: the scheduled increase to $3/$15 on 1 September 2026 will not occur. For teams that anchored budgets to the higher post-introductory figure, this is a straight 33% saving on the workhorse tier of Anthropic’s LLM API pricing.
Claude Haiku 4.5 and Sonnet 4.6
Claude Haiku 4.5, the small tier, costs $1 per million input and $5 per million output. Claude Sonnet 4.6, the previous workhorse, remains listed at $3/$15 — which now makes it dearer than its successor Sonnet 5, an unusual inversion worth catching in any LLM API pricing review of your model list.
Claude Fable 5: top-tier LLM API pricing
Claude Fable 5, Anthropic’s top tier, costs $10 per million input and $50 per million output tokens. Its caching economics are spelled out precisely: cache writes cost $12.50 per million for the 5-minute tier and $20 for the 1-hour tier, while cache hits cost just $1 per million — a tenth of the base input rate.
Caching multipliers, batch and fast mode
Anthropic’s caching multipliers apply across the range: 5-minute cache writes at 1.25x base input, 1-hour writes at 2x, and cache reads at 0.1x — a 90% discount on cached input. The Batch API halves both meters: Opus 5 drops to $2.50/$12.50, Sonnet 5 to $1/$5 and Haiku 4.5 to $0.50/$2.50 per million. In the other direction, fast mode (a research preview on the Claude API only) prices Opus 5 and Opus 4.8 at $10/$50, and US-only inference via inference_geo applies a 1.1x multiplier on all token categories. Anthropic web search bills separately at $10 per 1,000 searches.
The tokenizer footnote
Two structural notes round out Claude’s list. Claude 4.6 and later include the full 1M-token context window at standard rates, with no long-context premium — unusual among the providers here. And Claude 4.7-and-later models tokenize roughly 30% heavier than Sonnet 4.6 and earlier, so effective per-request costs shift even where per-token LLM API pricing stays flat.
Google Gemini Rates and the Promotional Clock
Gemini 3.7 Flash: promotional LLM API pricing until 31 December
Gemini 3.7 Flash costs $0.75 input / $3.75 output per million tokens — but this is promotional LLM API pricing that runs only through 31 December 2026. Batch processing halves it to $0.375/$1.875, and context caching charges $0.075 per million tokens plus $0.50 per hour of storage. If your 2027 budget assumes the promo rate, diarise the expiry now.
Gemini 3.5 Flash and the Flash-Lite budget line
Gemini 3.5 Flash costs $1.50/$9.00 per million, falling to $0.75/$4.50 in batch. Below it, Google runs a dense budget ladder: Gemini 3.5 Flash-Lite at $0.30/$2.50 and Gemini 3.1 Flash-Lite at $0.25/$1.50. The Lite pair competes directly with gpt-5.6-luna and Mistral Small 4 at the price floor.
Gemini 3.1 Pro Preview’s tiered rates
Gemini 3.1 Pro Preview uses prompt-length tiers: $2.00 input / $12.00 output per million for prompts up to 200k tokens, rising to $4.00/$18.00 above 200k. Batch halves both tiers. Above-200k input at $4.00 runs to more than five times Gemini 3.7 Flash’s $0.75 promo rate — tier boundaries are where Gemini LLM API pricing surprises the unwary.
Omni Flash Preview and video output
Gemini Omni Flash Preview shows where multimodal billing is heading: $1.50 per million input tokens, $9.00 per million text output tokens — and $17.50 per million video output tokens. Video output at nearly double the text rate is a new line item to watch; our review of the model itself is on the hub.
The free tier
The Gemini API keeps a genuine free tier: free-of-charge input and output tokens for most models via Google AI Studio, at limited rate limits. For prototyping and low-volume internal tools, that makes Google the default zero-budget entry point among the seven providers in this LLM API pricing round-up.
Mistral Rates: Open-Weight Economics
Mistral Large 3 undercuts Medium 3.5
Mistral’s list contains a genuine oddity: the open-weight Mistral Large 3 costs $0.50 input / $1.50 output per million — cheaper than the proprietary Mistral Medium 3.5 at $1.50/$7.50. The flagship being 5x cheaper on output than the mid-tier is a pattern found nowhere else in current LLM API pricing, and it makes Large 3 one of the strongest price-performance picks in Europe.
Small 4, Codestral and Ministral 3 LLM API pricing
Mistral Small 4 costs $0.15/$0.60 per million. The specialist line goes lower still: Codestral, the coding model, at $0.30/$0.90, and the edge-focused Ministral 3 family at $0.10/$0.10 for the 3B size, $0.15/$0.15 for the 8B and $0.20/$0.20 for the 14B. Symmetric input/output rates on Ministral make cost forecasting unusually simple.
GLM 5.2 on La Plateforme
Mistral’s La Plateforme also hosts third-party models: GLM 5.2 is listed at $1.40 input / $4.40 output per million, with cached input at $0.14. Hosting a Chinese open-weight flagship on a European platform at a 90% cache discount is exactly the kind of cross-pollination that keeps LLM API pricing competitive.
Mistral’s discount stack
Mistral’s discounts mirror the majors: cached input tokens cut input cost by up to 90%, and the Batch API halves prices for high-volume work. One premium runs the other way — regional inference adds 10%. That is the same shape as Anthropic’s 1.1x inference_geo multiplier, and a sign that data-residency surcharges are becoming a standard line in LLM API pricing.
DeepSeek Rates: The Price Floor of 2026
DeepSeek V4-Flash: the cheapest tokens on any list
DeepSeek’s deepseek-v4-flash (DeepSeek-V4-Flash-0731) costs $0.14 per million input tokens on a cache miss and $0.28 per million output tokens. On a cache hit, input falls to $0.0028 per million — a 98% discount, the deepest cache cut in this guide. This is the floor of current LLM API pricing: the $0.28 output rate anchors the 640-fold spread cited throughout this article.
DeepSeek V4-Pro rates
The larger deepseek-v4-pro (DeepSeek-V4-Pro-0813) costs $0.435 per million input on a cache miss, $0.003625 on a cache hit, and $0.87 per million output. Even DeepSeek’s premium model, in other words, prices below every Western budget tier except gpt-5-nano and Mistral’s smallest models. Note the model naming: V4-Flash and V4-Pro are the current pair on the official list — there is no separate chat/reasoner split any more.
Off-peak billing arrives on 16 August
From 16:00 UTC on 16 August 2026, DeepSeek switches to peak/off-peak billing. Peak hours are 01:00–04:00 and 06:00–10:00 UTC; outside those windows, off-peak rates are 50% of peak rates for both models. Time-of-day discounts are new to mainstream LLM API pricing, and they reward exactly the batch-shaped workloads that can wait a few hours.
Why a 98% cache discount resets expectations
Pair the numbers and the implication is stark: a cache-heavy workload on V4-Flash pays $0.0028 per million for repeated input — effectively rounding-error money. Once off-peak halving lands on 16 August, DeepSeek will have stacked three multiplicative levers under an already-lowest sticker. Competitors’ LLM API pricing teams will be studying that stack closely.
Alibaba Qwen Rates — and a Sourcing Caveat
Qwen3.8-Max LLM API pricing
Qwen3.8-Max, released on 3 August 2026, is priced at $2 per million input tokens and $6 per million output tokens, with a 1M-token context window. Cached input costs $0.25 per million — an 8x discount (87.5% off) on the $2 standard rate. That places Alibaba’s flagship squarely against Grok 4.6, which shares the identical $2/$6 sticker.
Qwen3.5-Plus and Qwen3.5-Flash
On Alibaba Cloud Model Studio, Qwen3.5-Plus costs $0.40/$2.40 per million input/output tokens, rising to $0.50/$3.00 above 256K input tokens. Qwen3.5-Flash costs $0.10/$0.40 — matching gpt-5-nano’s output rate with double its input rate. The tiering at 256K rather than 200k is a small but real difference when comparing long-context LLM API pricing across providers.
Free quota and the Model Studio caveat
Most Qwen models on Model Studio’s International (Singapore) deployment include a free quota of 1 million tokens, valid for 90 days after activation. One transparency note: Alibaba’s official pricing page could not be independently fetched for this guide, so Qwen figures here are sourced from OpenRouter, TechRepublic and BenchLM. Treat Model Studio itself as the canonical source before committing spend.
xAI Grok Rates: Simple Tokens, Separate Tools
Grok 4.6 LLM API pricing
xAI’s flagship grok-4.6, released on 12 August 2026 with a 500k context window, costs $2.00 input / $6.00 output per million tokens, with cached input at $0.50 for prompts under 200k tokens. At or above 200k tokens, the whole request is billed at $4.00/$12.00 with cached input at $1.00 — the same doubling pattern as Gemini’s Pro tier, applied to the entire request.
Grok 4.3 and grok-build
grok-4.3 offers a 1M-token context window at $1.25 input / $2.50 output per million, with cached input at $0.20 below 200k tokens; at or above 200k both meters double to $2.50/$5.00. The coding-focused grok-build-0.1 costs $1.00/$2.00. On raw output rate, Grok 4.3 is the cheapest 1M-context model in this entire LLM API pricing guide.
Tool calls billed separately
xAI bills Web Search, X Search and Code Execution separately at $5.00 per 1,000 calls — so an agentic workload pays a per-action fee on top of tokens. xAI is also the one provider in this round-up without a 50% batch discount, which materially changes its effective position for asynchronous workloads.
Stacking Discounts on LLM API Pricing
The single most useful table in this guide may be this one: the discount stack, provider by provider, as published in August 2026.
| Provider | Cache read discount | Cache write cost | Batch discount | Other levers & premiums |
|---|---|---|---|---|
| OpenAI | typically 90% off input | automatic | 50% | — |
| Anthropic | 90% (0.1x input) | 1.25x (5-min) / 2x (1-hr) | 50% | fast mode premium; inference_geo 1.1x |
| $0.075/M (3.7 Flash) | $0.50/hr storage | 50% | 3.7 Flash promo ends 31 Dec 2026; free tier | |
| Mistral | up to 90% off input | automatic | 50% | regional inference +10% |
| DeepSeek | 98% off input | automatic | 50% | off-peak 50% from 16 Aug 2026 |
| Alibaba (Qwen) | 87.5% (8x) on 3.8-Max | automatic | 50% | free 1M-token quota, 90 days |
| xAI | cached input $0.50/$0.20 | automatic | none | tools $5/1,000 calls; 200k doubling |
Prompt caching compared
Caching is the deepest lever in LLM API pricing today. The chart below plots the published cache-read discounts stated in this article: DeepSeek’s 98%, the 90% tier at Anthropic, OpenAI and Mistral, and Qwen3.8-Max’s 8x (87.5%) discount.
Caching mechanics differ, though. Anthropic charges for writes — 1.25x input for the 5-minute cache, 2x for the 1-hour tier — so caching pays only when reads outnumber writes. Google charges storage by the hour ($0.50 on 3.7 Flash). OpenAI, Mistral and DeepSeek apply discounts automatically on repeated prefixes. Model your hit rate before assuming the headline discount.
Batch APIs: the universal 50% lever
Every provider in this guide except xAI now offers a 50% batch discount on asynchronous work. The published examples: Anthropic’s Opus 5 falls to $2.50/$12.50, Sonnet 5 to $1/$5 and Haiku 4.5 to $0.50/$2.50; OpenAI’s gpt-5.6-sol to $2.50/$15.00; Gemini 3.7 Flash to $0.375/$1.875 and 3.5 Flash to $0.75/$4.50. If a workload does not need an interactive response, running it at sticker price is simply a 2x overspend.
Off-peak windows and time-shifting
DeepSeek’s peak/off-peak split, live from 16:00 UTC on 16 August 2026, defines peak as 01:00–04:00 and 06:00–10:00 UTC and halves both models outside those hours. No other provider yet publishes time-of-day LLM API pricing, but the economics of GPU utilisation make it a natural experiment for others to copy. Overnight report generation, embedding refreshes and evaluation runs are the obvious candidates to shift.
Premiums that work against you
The stack cuts both ways. Anthropic’s inference_geo US-only routing multiplies all token categories by 1.1x, and its fast mode research preview prices Opus 5 and Opus 4.8 at $10/$50 — double the standard rate. Mistral’s regional inference adds 10%. xAI’s $5 per 1,000 tool calls and Anthropic’s $10 per 1,000 web searches sit outside token maths entirely. Budget for the premiums with the same care as the discounts.
Choosing the Best Rates for Your Workload
The budget tier is where LLM API pricing competition is fiercest — this table collects every model in this guide with output at or below $2.50 per million, sorted cheapest first.
| Model | Provider | Input $/M | Output $/M |
|---|---|---|---|
| Ministral 3 (3B) | Mistral | $0.10 | $0.10 |
| Ministral 3 (8B) | Mistral | $0.15 | $0.15 |
| Ministral 3 (14B) | Mistral | $0.20 | $0.20 |
| deepseek-v4-flash | DeepSeek | $0.14 | $0.28 |
| gpt-5-nano | OpenAI | $0.05 | $0.40 |
| Qwen3.5-Flash | Alibaba | $0.10 | $0.40 |
| Mistral Small 4 | Mistral | $0.15 | $0.60 |
| deepseek-v4-pro | DeepSeek | $0.435 | $0.87 |
| Codestral | Mistral | $0.30 | $0.90 |
| gpt-5.6-luna | OpenAI | $0.20 | $1.20 |
| gpt-5.4-nano | OpenAI | $0.20 | $1.25 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | |
| Mistral Large 3 | Mistral | $0.50 | $1.50 |
| gpt-5-mini | OpenAI | $0.25 | $2.00 |
| grok-build-0.1 | xAI | $1.00 | $2.00 |
| Qwen3.5-Plus | Alibaba | $0.40 | $2.40 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | |
| Grok 4.3 | xAI | $1.25 | $2.50 |
High-volume, cost-first workloads
For classification, extraction, moderation and templated generation at scale, the shortlist is DeepSeek V4-Flash ($0.14/$0.28 with 98% cache hits and off-peak halving from 16 August), gpt-5-nano ($0.05/$0.40) and Qwen3.5-Flash ($0.10/$0.40). At these rates the engineering time spent switching usually costs more than a month of tokens — which is exactly why this end of LLM API pricing is where providers fight hardest.
Balanced production workloads
For customer-facing assistants and RAG, the workhorse tier clusters tightly: Claude Sonnet 5 ($2/$10, now permanent), gpt-5.6-terra ($2/$12), Gemini 3.5 Flash ($1.50/$9) and Qwen3.8-Max ($2/$6) all sit within touching distance. Grok 4.6 matches Qwen’s $2/$6. Here quality-per-pound, latency and caching mechanics matter more than the sticker, because the stickers barely differ.
Premium reasoning and agent workloads
At the top, choose between Claude Opus 5 ($5/$25), gpt-5.6-sol ($5/$30), Claude Fable 5 ($10/$50) and — for the deepest pockets — gpt-5.5-pro ($30/$180). Since Opus 4.8 through 4.5 are priced identically to Opus 5, pinning an older Opus for stability costs nothing extra. Agentic stacks must also budget tool fees: $5 per 1,000 calls at xAI, $10 per 1,000 web searches at Anthropic.
Long-context LLM API pricing compared
If your prompts run to hundreds of thousands of tokens, the rankings reshuffle. Grok 4.3 offers a 1M-token window at $1.25/$2.50 — though both meters double at 200k. Qwen3.8-Max pairs its 1M window with a flat $2/$6. Gemini 3.1 Pro Preview steps up to $4/$18 above 200k. Anthropic’s answer is structural: Claude 4.6 and later include the full 1M-token window at standard rates with no premium, which makes long-context LLM API pricing far easier to forecast on Claude than anywhere else.
Budget output rates, charted
The chart below plots output rates at the price floor, using the figures from the budget table above — the whole bottom tier of LLM API pricing now fits under Gemini 3.1 Flash-Lite’s $1.50.
A Practical Playbook to Cut Your LLM API Bill
Design for cache hits first
Structure prompts so the stable material — system message, tool definitions, reference documents — forms a repeated prefix, and the variable user turn comes last. That is what turns a $5 input rate into $0.50 at OpenAI, $2 into $0.20 at Anthropic, or $0.14 into $0.0028 at DeepSeek. On cache-heavy traffic, prompt architecture moves LLM API pricing more than model choice does.
Batch everything that can wait
Audit your traffic for anything asynchronous: nightly summaries, embeddings refreshes, evaluation suites, backfills. Move it to the provider’s batch endpoint and the meter halves — Sonnet 5 at $1/$5, gpt-5.6-sol at $2.50/$15, Gemini 3.7 Flash at $0.375/$1.875. Most teams find a third or more of their volume never needed to be interactive.
Time-shift to off-peak windows
From 16 August, DeepSeek workloads scheduled outside 01:00–04:00 and 06:00–10:00 UTC bill at half of peak rates. If DeepSeek is in your stack, a one-line scheduler change captures the discount. Watch for other vendors copying the model — time-of-day rates reward exactly the flexibility that batch discounts already do.
Diarise the promotional end dates
Two dates belong in your finance calendar now: 31 December 2026, when Gemini 3.7 Flash’s $0.75/$3.75 promo lapses, and the non-event of 1 September 2026, when Sonnet 5’s price will no longer rise to $3/$15. Promotional and permanent rates look identical in an invoice — a disciplined AI strategy tracks which is which before renewal negotiations.
Right-size models and route traffic
The 25x spread inside OpenAI’s GPT-5.6 family — sol at $5/$30 down to luna at $0.20/$1.20 — exists precisely so you can route easy requests to cheap models. Classify request difficulty, route accordingly, and reserve premium tokens for the queries that earn them. Wiring that routing into production is standard intelligent automation work, and it compounds with every discount above.
Re-price quarterly
Four of the facts in this guide — Sonnet 5’s permanence, DeepSeek’s off-peak switch, Grok 4.6’s 12 August launch and Qwen3.8-Max’s 3 August launch — are days or weeks old. Any LLM API pricing decision older than a quarter deserves a re-check against the References below.
The Bottom Line for August 2026
What this snapshot actually says
Strip the detail away and three truths remain. First, LLM API pricing has stratified into stable tiers — a premium band from $25 to $180 output, a $2-input workhorse band, and a budget floor below $2.50 output — so cross-provider switching within a tier is cheaper than ever. Second, the discount stack now matters more than the sticker: caching cuts input by 90–98%, batch halves everything at six of the seven providers, and off-peak billing arrives at DeepSeek on 16 August.
The one habit to build
Third, and most practically: dates now drive LLM API pricing as much as engineering does. Sonnet 5’s permanence, Gemini 3.7 Flash’s 31 December promo expiry and DeepSeek’s 16 August switchover all landed within weeks of each other. Put a quarterly re-pricing review in the calendar, keep the References below bookmarked, and treat every published rate — including every rate in this guide — as valid only until the next announcement.
FAQ: LLM API Pricing Questions Answered
What is LLM API pricing?
LLM API pricing is the pay-per-use rate charged for calling a large language model over an API, metered separately for input tokens (your prompt) and output tokens (the reply), and quoted per million tokens. It is distinct from consumer chat subscriptions, which charge per seat rather than per token.
Which provider has the cheapest LLM API pricing in 2026?
On sticker rates, DeepSeek: V4-Flash costs $0.14 input / $0.28 output per million, falling to $0.0028 input on cache hits, with a further 50% off-peak discount arriving on 16 August 2026. Among Western providers, gpt-5-nano ($0.05/$0.40) has the lowest input rate and Ministral 3 (3B) the lowest flat rate at $0.10/$0.10.
Which premium model offers the best value?
Claude Opus 5 at $5/$25 undercuts gpt-5.6-sol’s $5/$30 on output, and Anthropic prices Opus 4.8, 4.7, 4.6 and 4.5 identically. Claude Fable 5 ($10/$50) and gpt-5.5-pro ($30/$180) occupy the tier above. Value depends on your tokenizer overhead too — Claude 4.7+ produces roughly 30% more tokens for the same text.
How often does LLM API pricing change?
Constantly — August 2026 alone brought Sonnet 5’s rate becoming permanent, Grok 4.6’s launch, Qwen3.8-Max’s launch and DeepSeek’s off-peak announcement, with Gemini 3.7 Flash’s promo expiring 31 December. Treat any LLM API pricing table, including this one, as a dated snapshot and re-verify quarterly.
Does prompt caching really cut LLM API pricing by 90%?
Yes, on the input side: Anthropic bills cache reads at 0.1x base input, OpenAI’s cached input is typically 90% off, Mistral discounts up to 90%, and DeepSeek reaches 98%. But writes can cost extra (1.25x–2x at Anthropic; hourly storage at Google), so real savings depend on your cache hit rate.
Do all providers offer batch discounts?
All except xAI. The 50% batch discount is now effectively an industry standard — Anthropic, OpenAI, Google, Mistral, DeepSeek and Qwen all halve rates for asynchronous processing, per the price lists cited in this guide. xAI instead differentiates on simple flat rates plus per-call tool fees.
Is Qwen’s published LLM API pricing official?
Partially. The Qwen3.8-Max ($2/$6), Qwen3.5-Plus ($0.40/$2.40) and Qwen3.5-Flash ($0.10/$0.40) figures in this guide come from OpenRouter, TechRepublic and BenchLM because Alibaba’s own pricing page could not be independently fetched. Alibaba Cloud Model Studio remains the canonical source — verify there before committing spend.
How can you test LLM API pricing for free?
Two published routes exist. The Gemini API offers free-of-charge input and output tokens for most models via Google AI Studio, at limited rate limits. And most Qwen models on Model Studio’s International (Singapore) deployment include a free quota of 1 million tokens, valid for 90 days after activation. Both are enough to benchmark real workloads before spending.
Are long prompts charged extra?
Sometimes. Gemini 3.1 Pro Preview steps from $2/$12 to $4/$18 above 200k tokens; Grok 4.6 and 4.3 double for the whole request at 200k; Qwen3.5-Plus rises above 256K input. Anthropic is the exception: Claude 4.6+ includes the full 1M-token context window at standard rates.
References
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.