Token cost control has become the part of AI engineering that finance teams notice first. An InfoWorld feature published on 7 October 2026, “Five keys to controlling AI token costs” by contributing writer Matthew Tyson, names five levers: model routing, semantic caching, prompt caching, prompt discipline with reranking, and response constraints. Its argument is that large language model spend is now an architecture problem rather than a procurement one, because an application wrapped in agents and prompt templates “becomes a black box that consumes capital to spawn tokens.”

This guide takes those five keys and puts current numbers on them. Every price below comes from the vendors’ own pricing pages as they stood on 8 October 2026, and every saving in the charts is arithmetic on those prices, not a survey figure. One detail in the InfoWorld piece needs updating: OpenAI now charges a premium to write to its prompt cache on GPT-5.6 and later models, so the old split between “automatic and free” and “explicit and paid” caching no longer holds.

If you need to estimate a bill before you build, start with our AI token cost calculator. If the question is who owns the spend, our FinOps for AI guide covers budgets and chargeback. This article is about token cost control in the request path: what each key cuts, what it costs to run, where it backfires, and the order to apply the keys in.

Why Token Cost Control Is Now an Architecture Problem

token cost control five keys ai token costs b drill bit index stand with five graduated bits

InfoWorld’s opening comparison is with the cloud. Organisations spent a decade building FinOps teams to read cloud bills, then generative tools added a faster and more opaque layer of spend on top. Tyson calls the result “AI bill shock” and says the deeper problem is attribution: where exactly the money goes, and what value it delivers to the business.

Two properties of token pricing make that harder than the cloud version. First, cost scales with every request rather than with provisioned capacity, so a feature that takes off can multiply its bill overnight without any infrastructure change. Second, the meter runs on text the application itself generates, including prompts that grow as agents append tool results to their context. Token cost control therefore has to live in the code that builds prompts and parses answers, not in a monthly report.

Output tokens cost five times input

InfoWorld says output tokens are “almost universally priced 3x to 5x higher” than input tokens. On the flagship models listed in October 2026 the ratio is exactly 5x across the board, as the table below shows. That single fact explains why the length of a response matters as much as the length of a prompt, and why output belongs in every token cost control plan.

Model (October 2026)Input per 1M tokensOutput per 1M tokensCached read per 1MOutput vs input
Claude Fable 5.1$10.00$50.00$0.255x
Claude Opus 5.5$4.00$20.00$0.205x
Claude Sonnet 5.5$2.00$10.00$0.105x
Claude Haiku 5.5 (prompts up to 100k)$0.10$0.50$0.015x
GPT-6 Astra (up to 272k)$10.00$50.00$1.005x
GPT-6.1 Sol (up to 272k)$2.00$10.00$0.105x
GPT-6 Luna (up to 272k)$0.10$0.50$0.015x
Gemini 3.8 Flash (to 31 Dec 2026)$0.75$3.75$0.0755x

Token cost control starts with attribution

You cannot run token cost control on a bill you cannot split. The minimum is a per-request log of model, input tokens, cached tokens, output tokens and the feature or tenant that made the call. Every provider returns those counts in the usage block of each response, so the data already exists; many teams simply never store it. An AI gateway centralises the same telemetry across providers, which is why InfoWorld recommends one. Our guide to AI cost governance covers the budget controls that sit on top of that data.

Cheaper tokens do not mean a cheaper bill

Per-token prices keep falling. Claude Opus 5.5 lists at $4 per million input tokens, against $5 for Claude Opus 5, and Claude Haiku 5.5 launched on 7 October at $0.10 for prompts up to 100,000 tokens. Yet agentic workloads send far more tokens per task than chat ever did, so many bills keep rising. Our report on Meta’s tokenmaxxing leaderboard showed how quickly raw token volume becomes the number people chase. Token cost control is the discipline of keeping that volume tied to value.

The Five Keys to Token Cost Control at a Glance

token cost control five keys ai token costs c ice chest with a hinged lid and two side handles

Each key attacks a different part of the bill. Routing changes the price per token. The two caches either remove whole calls or reprice repeated tokens. Reranking shrinks what you send, and response constraints shrink what comes back. The table summarises the five, using only figures that the sources themselves state.

KeyWhat it cutsSaving the source statesMain risk
Model routingPrice per tokenRouteLLM: up to 85% lower cost at 95% of GPT-4 performance on MT BenchHard queries sent to a model that cannot answer them
Semantic cachingWhole calls for repeated intentsA hit costs no generation tokens at allGeneric or stale answers
Prompt cachingPrice of repeated prefix tokensReads at 0.025x to 0.1x of the input priceSilent misses when the prefix changes
Prompt discipline and rerankingInput tokens per requestInfoWorld: prompt tokens cut by 80% or moreDropping a chunk the answer needed
Response constraintsOutput tokensPreamble removed from the most expensive tokensStarving the model of room to reason

The keys stack. The worked example later in this article applies all five to one support assistant and takes a modelled monthly bill from $62,800 to about $7,213. The order matters almost as much as the levers themselves, because some token cost control measures are free to switch on while others need evaluation work before they are safe to ship.

Price levers and volume levers in token cost control

It helps to sort the keys into two groups. Prompt caching and model routing are price levers: the model sees the same tokens, or the same kind of tokens, but each one costs less. Semantic caching, reranking and response constraints are volume levers: fewer tokens are sent or generated at all. Price levers are usually faster to ship, while volume levers usually save more on retrieval-heavy applications.

Key One: Model Routing Puts Token Cost Control in the Request Path

token cost control five keys ai token costs d chafing dish with a domed lid on a frame stand

“Don’t use a nail gun if a thumbtack will suffice,” Tyson writes. Teams prototype on the most capable model available so they are not fighting model limits while they work out requirements. That is sensible in development and expensive in production, where the right question is which is the least capable model that still meets the quality bar for each kind of request.

What one routing decision is worth

The chart shows the spread at list price. It prices one million requests, each with 2,000 input tokens and 400 output tokens, on each model’s standard rate. For identical traffic, the most capable model costs 100 times more than the cheapest one.

List-price cost of 1 million requests (2,000 input + 400 output tokens each)
Claude Fable 5.1 $40,000
Claude Opus 5.5 $16,000
Claude Sonnet 5.5 or GPT-6.1 Sol $8,000
Gemini 3.8 Flash $3,000
Haiku 5.5 or GPT-6 Luna $400

Tokenizers differ between models, so the same text produces different token counts on different models, even from one vendor. Count tokens with each provider’s own counter before you trust a cross-model comparison like this one.

Routers: RouteLLM and Semantic Router

Manual tuning works for a handful of endpoints. Larger estates need a routing layer. RouteLLM, the open-source framework from the LMSYS team, ships trained routers that its maintainers say “reduce costs by up to 85% while maintaining 95% GPT-4 performance” on benchmarks such as MT Bench. Each request carries a cost threshold that you calibrate on your own traffic, which sets how often the strong model gets called.

Semantic Router from Aurelio Labs takes a different approach. It classifies intent by embedding similarity against example utterances, so simple classification, text parsing and intent detection can go to a small model without spending a model call to decide. Both tools help with token cost control, but neither replaces an evaluation set: a routing threshold is only as good as the queries it was calibrated on.

Gateways, cascades and per-tenant budgets

InfoWorld recommends wrapping routing logic in an AI API gateway, naming Kong, Cloudflare AI Gateway and Portkey as examples. A mature gateway can run “cascade routing”, falling back to a cheaper or open-source model when the primary hits rate limits or latency spikes, and it can enforce hard token budgets per tenant or per microservice. Tyson calls it “a load balancer for cognitive tasks.” Gateways add their own cost and complexity, so as a token cost control tool they earn their place once you have several providers or many internal consumers.

When routing backfires

Routing has two hidden costs. Prompt caches are scoped to a model, so traffic split across three models keeps three separate caches warm, and each one gets fewer hits. And a cheap model that needs two retries or a longer conversation to finish a task is not cheap at all. Measure cost per completed task, not cost per request. Before building a multi-model cascade, test the simpler option of running the most capable model at a lower effort setting on the same traffic.

Running models on your own hardware

InfoWorld also notes a trend it calls “repatriation of compute”: running open-weight models on local hardware instead of renting endpoints. That swaps a per-token bill for capacity planning, power and hardware refresh cycles. It suits steady, high-volume, low-complexity traffic, and it is the logical end point of token cost control for workloads that never need a frontier model.

Key Two: Semantic Caching as Token Cost Control for Repeat Questions

token cost control five keys ai token costs e kitchen colander with two loop handles

Conventional caching maps an exact key to a stored value. Natural language breaks that. “How do I reset my password?” and “I forgot my login info” are different strings with identical intent, as InfoWorld points out, so an exact-match cache misses both and every rephrasing pays the full token price.

How a semantic cache works

A semantic cache embeds each incoming prompt, searches previously answered prompts by vector similarity, and returns the stored answer when the similarity clears a threshold. Tyson’s example threshold is 0.92. A hit spends no generation tokens and returns in milliseconds rather than seconds. GPTCache from Zilliz is a purpose-built option whose project page promises to “slash your LLM API costs by 10x”. Teams already running PostgreSQL can store and match the embeddings with the pgvector extension instead of adding a new service.

What a cache lookup costs

The lookup is not free, but it is close. OpenAI lists text-embedding-3-small at $0.02 per million tokens, so embedding a 200-token question costs $0.000004 and a million lookups cost $4. The vector search runs on a database you already pay for. Against a generation call that costs between $0.0004 and $0.04 in the routing chart above, the lookup is a rounding error, which is why semantic caching can be the strongest token cost control lever on repetitive traffic.

Threshold tuning and semantic flattening

The risk is what InfoWorld calls semantic flattening. Set the threshold too low and the application starts serving generic, recycled answers to nuanced questions. Set it too high and the hit rate collapses towards zero. Tune it on logged real queries, check a sample of hits by hand every week, and expire entries when the underlying documents change, or the cache will keep answering from last quarter’s policy page.

Where semantic caching pays, and where it does not

InfoWorld is blunt that for open-ended, creative applications semantic caching is “practically useless”. It shines in retrieval-augmented generation, customer support bots and internal knowledge bases, where “users ask the same 20 questions a thousand different ways.” Anything personalised needs care. A cached answer that includes one customer’s account details must never be served to another customer, so scope cache keys by tenant, or skip caching for personal data entirely.

Exact-match gateway caches are not semantic caches

Check what your gateway actually caches. Cloudflare’s AI Gateway documentation says its caching “applies only to identical requests”, and that semantic search for caching is planned for the future. Exact-match caching still helps with retries and repeated scheduled jobs, but it will not catch a paraphrase. Treat it as a complement to a semantic layer in your token cost control stack, not a substitute for one.

Key Three: Prompt Caching, the Cheapest Token Cost Control

token cost control five keys ai token costs f caulking gun with a cartridge and trigger handle

Semantic caching stores answers. Prompt caching stores the processed prefix of a prompt (the system instructions, tool definitions, documents and conversation history that repeat from call to call) so the provider does not have to recompute it. InfoWorld puts the discount on cached tokens at 50% to 90%. Current price lists go further: on the newest flagship models, reads cost between 0.025x and 0.1x of the normal input price.

ProviderHow it is switched onCache writeCache readMinimum prefixLifetime
Anthropic (Claude)One top-level cache_control field, or up to 4 explicit breakpoints1.25x input (5 minutes) or 2x (1 hour)0.1x; 0.05x on Opus 5.5 and Sonnet 5.5; 0.025x on Fable 5.1512 tokens on current models5 minutes, refreshed free on each hit; 1-hour option
OpenAI (GPT-5.6 and later)Implicit by default; explicit breakpoints optional1.25x input0.1x; 0.05x on GPT-6.1 Sol1,024 tokens30 minutes after the most recent use
Google (Gemini API)Implicit on Gemini 2.5 and newer; explicit cache objects for controlExplicit storage: $0.50 per 1M tokens per hour on Gemini 3.8 Flash0.1x on Gemini 3.8 Flash4,096 tokens on Gemini 3.xSet by you on explicit caches

Anthropic: breakpoints and two lifetimes

Anthropic’s caching is opt-in. You either add one cache_control field at the top level of the request, which caches everything up to the last cacheable block and moves forward as a conversation grows, or you place up to four breakpoints yourself. The default lifetime is five minutes, refreshed at no cost each time the entry is read, and a one-hour lifetime costs 2x the input price to write. The pricing page says caching “pays off after one cache read” on the five-minute option. Claude Sonnet 5.5 cache reads were halved on 7 October; see our report on the Claude Haiku 5.5 launch and the Sonnet 5.5 cache cut.

OpenAI now charges to write the cache

InfoWorld describes OpenAI’s caching as automatic, in contrast to explicit providers where you pay “a write premium” up front. That contrast no longer holds for OpenAI’s newest models. Its prompt caching guide says that for GPT-5.6 and later, “cache writes cost 1.25× the standard, uncached input-token rate”. Reads cost 0.1x, or 0.05x on GPT-6.1 Sol, and developers can now place explicit breakpoints, with up to four cache writes per request. Entries stay eligible for 30 minutes after their most recent use. The 1,024-token minimum and the exact-prefix rule remain.

Gemini: implicit by default, explicit for control

Google’s Gemini API turns on implicit caching by default for Gemini 2.5 and newer models, with a minimum of 4,096 input tokens on the Gemini 3.x models it lists. Explicit cache objects let you decide what stays cached and for how long, and they are billed for storage: $0.50 per million tokens per hour on Gemini 3.8 Flash through 31 December 2026, rising to $1.00 from 1 January 2027. Cached reads on that model cost $0.075 per million tokens, against $0.75 for fresh input. For Gemini, token cost control starts with checking that a prompt clears the 4,096-token minimum before expecting any discount.

Prefix discipline: the silent cache killers

Every provider matches on an exact prefix. Change one byte early in the prompt and everything after it misses. InfoWorld’s example is a timestamp or user ID at the top of the system instructions, which will “instantly invalidate the cache”. Other common culprits are tool lists assembled in a different order on each call, JSON serialised with unsorted keys, and request IDs in the system prompt. Put stable content first and volatile content last, then read the cached-token count in every response to confirm hits are happening.

Break-even arithmetic

The chart prices ten requests that share a 50,000-token prefix on Claude Sonnet 5.5 at its October list prices: $2 per million input tokens, $2.50 to write a five-minute entry, $4 to write a one-hour entry, and $0.10 to read. One write plus nine reads costs 17 cents against a dollar uncached, an 83% saving. GPT-6.1 Sol carries the same four prices, so the same arithmetic applies there. This is the cheapest token cost control there is, because the model sees exactly the same tokens.

Ten requests sharing a 50,000-token prefix on Claude Sonnet 5.5
No caching (500,000 tokens at $2 per million) $1.000
1-hour cache (one write at $4, nine reads at $0.10) $0.245
5-minute cache (one write at $2.50, nine reads at $0.10) $0.170

Key Four: Reranking and Prompt Discipline for Input Token Cost Control

“Just because a model can consume gargantuan amounts of tokens, doesn’t mean it should,” Tyson writes. With context windows of a million tokens or more, it is tempting to paste a whole PDF, codebase or log archive into the prompt and let the model sort it out. Every one of those tokens is billed, every time.

Long context is not free context

At list price, a single prompt that fills a one-million-token window costs $4 on Claude Opus 5.5 and $2 on Claude Sonnet 5.5 before the model writes a word. On GPT-6.1 Sol, prompts above 272,000 tokens move to a long-context rate of $4 per million input tokens, double the short-context price. A team sending 500 such prompts a day on Opus 5.5 spends $2,000 a day on input alone. Our comparison of RAG, fine-tuning and long context covers when a large window really is the right tool.

Lost in the middle and context rot

Bigger prompts can also be worse prompts. The Stanford-led 2023 paper “Lost in the Middle” found that performance “is often highest when relevant information occurs at the beginning or end of the input context”, and that it degrades when the answer sits in the middle. Chroma’s July 2025 “Context Rot” report tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3, and found that performance “grows increasingly unreliable as input length grows”, even on simple tasks. Token cost control and answer quality point the same way here.

Retrieve wide, rerank narrow

The fix InfoWorld prescribes is “a strict RAG diet”. Naive retrieval pulls the top 20 chunks from a vector database and concatenates them into the prompt. A better pipeline casts a wide net of around 50 candidates from cheap vector search, then passes them to a small cross-encoder reranker, such as Cohere Rerank or one of BAAI’s open-source BGE rerankers. The reranker scores each chunk against the actual question, and only the top three or four reach the expensive model.

What reranking saves and what it costs

InfoWorld says a cross-encoder might add about 100 milliseconds of latency but “routinely slashes prompt token counts by 80% or more”. The chart shows what an 80% cut in retrieved context means for one million requests on Claude Opus 5.5, with 500-token chunks.

Retrieved-context input cost per 1 million requests on Claude Opus 5.5 ($4 per million input tokens)
Top 20 chunks, 10,000 tokens per request $40,000
Top 4 reranked chunks, 2,000 tokens per request $8,000

The reranker has its own bill. Cohere prices hosted Rerank by the “search”, defined as one query with up to 100 documents, and documents over 500 tokens are split into extra chunks that count separately. For dedicated capacity, its Model Vault lists a medium Rerank 4 Fast instance at $5 an hour or $3,250 a month. An open-source BGE reranker on your own GPU swaps that for hosting costs. At this volume, either option costs far less than the $32,000 of input it removes.

Tighten the prompt itself

Reranking trims retrieved context; prompt discipline trims everything else. Remove instructions the model follows anyway, delete few-shot examples that no longer change outputs, summarise long conversation history instead of resending it, and trim verbose tool results before they go back into context. Agent loops gain the most from this kind of token cost control, because each tool result is re-sent on every later turn. Our guide to preventing agent cost overruns covers loop limits.

Key Five: Response Constraints as Output Token Cost Control

Models are trained to be polite. Ask for data and you may get “Certainly! Here is the JSON you requested” before it and an offer of further help after it. In a chat window that is charming. In an API pipeline it is, in Tyson’s phrase, “a leak in the finops hull”, because output is the most expensive token type on every model in the price table.

What verbosity costs

Claude Opus 5.5 charges $20 per million output tokens. An extra 200 tokens of preamble and sign-off on one million responses a month is 200 million tokens, or $4,000 a month, for text that no user or parser needs. Output-side token cost control starts with measuring average response length per endpoint and asking how much of it is doing real work.

max_tokens is a circuit breaker, not a formatter

InfoWorld’s advice is to treat max_tokens “as a hard circuit breaker to prevent runaway generation loops”, not as a way to shorten answers. Relying on the cap to keep answers short tends to produce truncated, unparseable JSON, which then costs a retry. Set the cap comfortably above the longest legitimate answer, alert whenever responses hit it, and control length with instructions and schemas instead.

Stop sequences end generation early

Stop sequences tell the model to stop the moment a marker appears. If an endpoint only needs a SQL query, a stop sequence on the closing code fence, or on a word such as “Explanation:”, ends the call as soon as the useful output is complete. The major APIs all support them, which makes them the cheapest output-side token cost control to add.

Structured outputs remove the chatter

Structured outputs constrain a response to a JSON schema you supply. OpenAI’s structured outputs guide says the feature “ensures the model will always generate responses that adhere to your supplied JSON Schema”, and Anthropic offers the same through an output format setting and strict tool definitions. A schema removes preamble by construction, and it also removes the retries you would otherwise pay for when a response fails to parse. It is one of the few token cost control measures that improves reliability at the same time.

Leave room to reason

The balance, as InfoWorld puts it, is that models “think” by generating tokens. Force a model to answer a hard logic problem with a single true or false and accuracy falls. The goal is not the fewest words; it is that every output token does useful work. On current reasoning models much of that thinking happens in tokens billed as output, so the effort setting is the real lever. Claude Opus 5.5 defaults to medium effort, and lower effort means fewer thinking tokens on routine requests.

Two Token Cost Control Levers the Five Keys Leave Out

InfoWorld’s five keys all act on live requests. Two more levers act on how and when you send them, and both appear on the providers’ own price lists.

Batch processing halves the price

Anthropic’s Batch API processes requests asynchronously “with a 50% discount on both input and output tokens”, and the discount stacks with prompt caching. OpenAI’s batch tier is half price too: GPT-6.1 Sol drops from $2 and $10 to $1 and $5 per million input and output tokens. Google lists a 50% batch reduction for Gemini. Overnight document processing, evaluation runs, embedding back-fills and report generation rarely need an answer in seconds, so they should rarely pay real-time prices. Moving them to batch is token cost control with no quality trade-off at all.

Effort and thinking budgets

Reasoning models spend output tokens thinking before they answer. Anthropic’s current models expose an effort setting that runs from low to max. Low suits classification and simple chat, while hard coding and agent work earn the higher settings. Per-route effort is one of the simplest token cost control settings to change, but measure quality on a sample of real requests before lowering a default, and judge the result by cost per completed task.

A Worked Token Cost Control Example: One Support Bot

To see how the keys stack, take a hypothetical customer support assistant that handles one million questions a month. Every number below is arithmetic on the October 2026 list prices quoted above. The traffic shape, hit rates and routing share are illustrative assumptions, stated so you can replace them with your own.

The baseline

Each request sends a 3,000-token system prompt with tool definitions, 10,000 tokens of retrieved context (the top 20 chunks of 500 tokens) and a 200-token question, all to Claude Opus 5.5. Answers average 500 tokens. Input comes to 13,200 million tokens at $4, or $52,800 a month, and output to 500 million tokens at $20, or $10,000. The total is $62,800 a month before any token cost control.

Applying the five token cost control keys in order

StepChangeAssumptionMonthly bill
BaselineAll traffic on Claude Opus 5.5, no caching13,200 input and 500 output tokens per request$62,800
1. Prompt cachingCache the 3,000-token prefixRewritten at most every 5 minutes: 8,640 writes in a 30-day month$51,524
2. RerankingTop 4 of 50 candidates instead of top 20Retrieved context falls to 2,000 tokens; one $3,250-a-month reranker instance$22,774
3. Response constraintsSchema output, no pleasantriesAnswers fall from 500 to 300 tokens$18,774
4. Semantic cachingServe repeated intents from cache20% of questions hit; embeddings cost $4$15,674
5. Model routingSend simple questions to Haiku 5.570% of remaining calls routed down, each model with its own cache$7,213

What the token cost control example shows

The modelled bill falls by about 89%. The largest single step is reranking, because retrieved context was the biggest block of input. Prompt caching saves $11,276 on the prefix alone at almost no engineering cost, even on the pessimistic assumption that the entry is rewritten every five minutes. By the last step the fixed $3,250 reranker instance makes up 45% of what remains, which is a sign the variable token bill is finally under control.

Modelled monthly bill after each key (support bot, 1 million questions)
Baseline $62,800
After prompt caching $51,524
After reranking $22,774
After response constraints $18,774
After semantic caching $15,674
After model routing $7,213

Your numbers will differ, because tokenizers, hit rates and routing shares vary by application. What transfers is the method: price each bucket of tokens separately, apply one key at a time, and re-measure after each. A spreadsheet with one row per bucket (prefix, retrieved context, question, output) is enough to plan token cost control for most applications.

The Right Order for Token Cost Control Work

The keys are not equally risky, so do not apply them in the order InfoWorld lists them. Start with token cost control changes that cannot affect answer quality, then move on to the ones that need an evaluation set.

Measure before any token cost control change

Log tokens by model, endpoint and tenant before changing anything. Without a baseline you cannot prove a saving, and you cannot find the endpoint that drives the bill. One verbose endpoint or one runaway agent loop can dominate a month’s spend, and only per-endpoint logs will show it.

Free wins: caching and batching

Prompt caching and batch processing change the price, not what the model reads or writes, so they are safe to ship first. Fix the prefix order, add cache breakpoints, confirm the cached-token counts, and move every non-urgent job to batch. These two steps are the lowest-risk token cost control work you can do.

Changes that need an evaluation set

Response constraints, reranking, semantic caching and routing all change what the model sees or which model answers. Each needs a set of real questions with known good answers, run before and after the change. Routing comes last because it is the hardest to get right, and because caching and trimmed prompts change the economics it is trying to optimise.

Token Cost Control Mistakes That Quietly Raise the Bill

Most overspend does not come from choosing the wrong model. It comes from small implementation details that defeat the token cost control measures already in place.

Caching a prefix that changes

A timestamp, request ID or reshuffled tool list near the top of the prompt turns every cache write into a miss. You pay the write premium on Anthropic and on OpenAI’s newer models and never collect the discount.

Treating max_tokens as a length control

Truncated JSON fails to parse, the request is retried, and you pay for the output twice. Use schemas and instructions for length, and keep the cap as a safety limit.

Routing on price alone

A cheaper model that fails more often costs more per resolved request. Route by measured quality on each class of query, and re-check after every model update.

Letting a semantic cache go stale

Cached answers outlive the documents they came from. Tie cache entries to source versions or set expiry times, and never share cached personal answers across users.

Optimising requests, not tasks

Agents can turn one user request into dozens of model calls. Token cost control has to cap loops and tool calls per task, not just tokens per call.

Token Cost Control FAQ

What is token cost control?

Token cost control is the set of engineering practices that keep language model spend in proportion to the value an application delivers. It covers choosing a model per request, caching answers and prompts, trimming input and constraining output, plus the logging that shows which of those is working.

Which token cost control key saves the most?

It depends on where your tokens go. Applications with large retrieved contexts save most from reranking. Applications with long, stable system prompts save most from prompt caching. High-volume, simple traffic saves most from routing. Price each bucket of tokens separately to find out which applies to you.

Does prompt caching work with batch processing?

Yes. Anthropic’s pricing documentation says its caching multipliers “stack with other pricing modifiers, including the Batch API discount”. Check each provider’s batch terms for the specific models you use.

Is semantic caching safe for customer data?

Only with care. Scope cache keys by tenant or user, never serve one customer’s personalised answer to another, and expire entries when policies or source documents change.

How do I know my prompt cache is working?

Every provider reports cached tokens in the usage data of each response. If that number stays at zero across requests that share a prefix, something early in the prompt is changing on every call.

References