GPT-6.1 Sol Ultrafast is now rolling out across OpenAI’s API, Codex and ChatGPT Work. The new speed tier runs OpenAI’s mid-priced GPT-6.1 Sol model at “up to 8x faster speeds than Sol Standard”, the company’s developer account announced on X at 18:26 UTC on Thursday 8 October 2026, promising “near-Astra intelligence” for people who want to “build as fast as the ideas come”.
The speed comes at a price. In the API, GPT-6.1 Sol Ultrafast costs $12 per million input tokens and $60 per million output tokens, six times the Standard rate, a figure TestingCatalog flagged within minutes and OpenAI’s pricing page confirms. In Codex and ChatGPT Work, it is limited to the $500 Pro plan and eligible Enterprise and Edu workspaces, and it burns through included usage eight times faster.
This article explains what OpenAI launched, how fast GPT-6.1 Sol Ultrafast really is, what it costs in the API and in ChatGPT plans, how to switch it on, which hardware may be behind it, and when the premium is worth paying. For background on the model itself, see our report on the launch of GPT-6.1 Sol at DevDay.
Table of contents
- What OpenAI Launched With GPT-6.1 Sol Ultrafast
- How Fast Is GPT-6.1 Sol Ultrafast?
- GPT-6.1 Sol Ultrafast Pricing in the API
- GPT-6.1 Sol Ultrafast in Codex and ChatGPT Work
- How to Use GPT-6.1 Sol Ultrafast in the API
- Who Is Serving GPT-6.1 Sol Ultrafast?
- When GPT-6.1 Sol Ultrafast Is Worth the Premium
- What GPT-6.1 Sol Ultrafast Means for the Market
- Frequently Asked Questions About GPT-6.1 Sol Ultrafast
- References
What OpenAI Launched With GPT-6.1 Sol Ultrafast
GPT-6.1 Sol arrived at OpenAI’s DevDay on 29 September as a cheaper alternative to the flagship GPT-6 Astra. OpenAI’s model page describes it as delivering “near-Astra performance at a lower cost for complex coding, computer use, and professional work”. Ultrafast is not a new model. It is a faster way of serving the same model, which OpenAI calls a service tier.
The OpenAI API changelog entry for 8 October says the company “added Ultrafast mode for GPT-6.1 Sol in the Responses API”, designed “to reduce the time between generated output tokens”. It adds that the tier “is available to all API users, subject to rate limits, with global processing and US and EU data residency”.
The GPT-6.1 Sol Ultrafast launch at a glance
| Item | GPT-6.1 Sol Ultrafast | Source |
|---|---|---|
| Released | Thursday 8 October 2026 | API and Codex changelogs |
| Where | API (Responses API), Codex, ChatGPT Work | OpenAI Developers on X |
| Speed claim | Up to 8x faster than Sol Standard | OpenAI Developers on X |
| API price, prompts up to 272K tokens | $12 input, $0.60 cached input, $15 cache writes, $60 output per 1M tokens | OpenAI API pricing |
| Multiple of Standard | 6x in the API; 8x of included usage in ChatGPT plans | Model page, ChatGPT pricing docs |
| ChatGPT plans | Pro $500, eligible Enterprise and Edu | Codex changelog |
| Data residency | US and Europe (EEA plus Switzerland) | Codex speed docs, API changelog |
What the launch video shows
OpenAI’s 14-second clip is deliberately simple. White text types out “This is GPT-6.1 Sol in Standard mode”, then the same sentence in Ultrafast mode appears almost at once, followed by “Try it in Codex, ChatGPT Work, and the API today.” It is a demonstration of perceived speed, not a benchmark, so the rest of this article looks at what OpenAI’s documentation and independent testers actually measured.
How Fast Is GPT-6.1 Sol Ultrafast?
OpenAI’s headline is “up to 8x faster speeds than Sol Standard”. The word “up to” matters. OpenAI’s documentation describes Ultrafast as reducing “the time between generated output tokens”, which is generation speed once the model starts writing. It does not promise a faster first token, faster tool calls or a faster finished task.
The written docs are also more careful than the tweet. OpenAI’s Codex speed page states an 8x figure only for GPT-6 Astra: “GPT-6 Astra Ultrafast generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex.” At DevDay, OpenAI’s own materials put Ultrafast at roughly 300 tokens per second and “up to 6x” faster in the API, as reported at the time.
An early independent test of GPT-6.1 Sol Ultrafast
The only public measurement so far comes from Kingy AI, which ran a small matched pilot on launch day and published it as “Is the Time Saved Worth the Price?”. It ran four matched pairs on each of two small tasks, using the same prompt, the same reasoning effort and the same client, with only the service tier changed.
The chart shows Kingy’s median completion times in seconds. GPT-6.1 Sol Ultrafast was about 3.7 times quicker on a small coding fix and about 3.9 times quicker on a research brief.
Quality did not change. Every coding answer passed all 30 of Kingy’s checks on both tiers, and both tiers missed the same condition in the research rubric. Kingy estimated the faster tier cost about 2.1 cents more per coding answer. With only 16 first attempts, treat this as a useful early signal rather than a verdict.
Why a whole task speeds up less than the tokens
A coding agent spends time reading files, calling tools, running tests and waiting on networks, and none of that gets faster when token generation does. Suppose a task takes 100 seconds, of which 60 are spent generating text. If that 60 seconds becomes eight times faster, it shrinks to 7.5 seconds, and the task takes 47.5 seconds overall: about 2.1 times faster, not eight.
That is why GPT-6.1 Sol Ultrafast helps most in tight, generation-heavy loops, such as an engineer iterating with Codex on a small change, and least in long agent runs dominated by test suites or slow websites.
GPT-6.1 Sol Ultrafast Pricing in the API
OpenAI’s API pricing page lists every processing tier for GPT-6.1 Sol. Short context means prompts of up to 272,000 input tokens; longer prompts are charged at higher long-context rates for the whole request.
Every GPT-6.1 Sol tier side by side
| Tier (per 1M tokens) | Input | Cached input | Cache writes | Output | Long-context output |
|---|---|---|---|---|---|
| Batch and Flex | $1.00 | $0.05 | $1.25 | $5.00 | $7.50 |
| Standard | $2.00 | $0.10 | $2.50 | $10.00 | $15.00 |
| Fast (up to 2.5x speed) | $4.00 | $0.20 | $5.00 | $20.00 | $30.00 |
| Ultrafast (up to 8x speed) | $12.00 | $0.60 | $15.00 | $60.00 | $90.00 |
The model page states the rule simply: “Ultrafast mode prices are 6x Standard.” Cached input keeps its discount at every tier, and regional processing for data residency adds 10% on top.
GPT-6.1 Sol Ultrafast against GPT-6 Astra
The most interesting comparison is with OpenAI’s flagship. GPT-6 Astra costs $10 per million input tokens and $50 per million output on Standard. GPT-6.1 Sol Ultrafast costs $12 and $60, only 20% more per token than Astra at normal speed. Astra’s own Ultrafast tier costs $60 and $300, five times the price of GPT-6.1 Sol Ultrafast.
The chart compares output prices per million tokens. For a team that finds Sol’s quality good enough, Ultrafast buys a large speed gain for roughly what Astra costs at Standard speed.
What one coding turn costs on GPT-6.1 Sol Ultrafast
Take a coding turn that sends 20,000 fresh input tokens and gets 2,000 tokens back. On Standard it costs $0.04 for input plus $0.02 for output, or $0.06. On GPT-6.1 Sol Ultrafast it costs $0.24 plus $0.12, or $0.36. A team running 1,000 such turns a day would pay $60 a day on Standard and $360 on Ultrafast, or about $1,320 against $7,920 over 22 working days.
Caching narrows the absolute gap but not the ratio. If 15,000 of those 20,000 input tokens are cached, the Standard turn falls to about $0.032 and the Ultrafast turn to about $0.189. The multiple stays close to six because every line of the price list is scaled by the same factor.
When the time saved pays for itself
Kingy’s pilot saved about seven seconds per coding answer for about 2.1 cents. If an engineer’s time costs the business $60 an hour, seven seconds is worth about 12 cents, so the faster answer pays for itself several times over, as long as someone is actually waiting for it. For an unattended overnight job, nobody is waiting, and the premium buys nothing.
GPT-6.1 Sol Ultrafast in Codex and ChatGPT Work
Inside OpenAI’s own apps, access is narrower and the billing works differently. The ChatGPT and Codex changelog for 8 October says: “Ultrafast is available on Pro $500 and eligible Enterprise and Edu plans. Enterprise access is off by default; workspace owners must enable it.”
ChatGPT Work and Codex share one usage pool, so time spent on GPT-6.1 Sol Ultrafast in either product draws on the same allowance. Our earlier analysis of ChatGPT’s $500 Pro tier explains why OpenAI built that plan around speed.
Who can use GPT-6.1 Sol Ultrafast in ChatGPT
| Plan | Ultrafast access | How it is charged |
|---|---|---|
| Plus, Pro $100, Pro $200 | No, “even with purchased credits” | Not applicable |
| Pro $500 | Yes | Included usage at 8x the Standard rate, then credits at 6x |
| Enterprise, credit or usage-based | Yes, off by default; owners enable per user or workspace | Credits or pay-as-you-go at 6x, under existing spend controls |
| Legacy Enterprise on rate limits | No | Not supported |
| Eligible Edu | Yes | Credits at 6x |
| Codex with an API key | Yes | API token prices; ChatGPT multipliers do not apply |
The 8x allowance trap
The two multipliers are easy to confuse. In the API you pay six times the Standard token price. Inside a ChatGPT plan, GPT-6.1 Sol Ultrafast uses your included allowance at eight times the Standard rate, according to OpenAI’s ChatGPT Work and Codex pricing docs. Fast mode, by comparison, uses it at 2.5 times.
At DevDay, OpenAI described the $500 Pro plan as carrying 25 times the Plus allowance. Spent entirely on GPT-6.1 Sol Ultrafast, that shrinks to roughly three times the Plus allowance at Standard speed (25 divided by 8). Heavy users should expect to reach the credits stage sooner than they might assume.
Enterprise controls
For Enterprise workspaces, Ultrafast is off until an owner enables it for selected users or the whole workspace through workspace permissions. Existing per-user spend controls then apply. That is sensible: a handful of engineers running GPT-6.1 Sol Ultrafast all day could move a monthly bill noticeably, so start with a named pilot group. Treat the switch like any other cybersecurity or spending control, with named users, a review date and a record of who enabled it.
How to Use GPT-6.1 Sol Ultrafast in the API
Switching tiers is a one-line change. In the Responses API, set the model to gpt-6.1-sol and the service_tier parameter to ultrafast. Nothing else about the request changes, and the response reports the tier that served it. OpenAI’s Ultrafast mode guide covers the setup.
Use WebSockets for agents
OpenAI says it “strongly” recommends WebSocket connections, “especially for agentic applications that make many tool calls in quick succession”. Without a persistent connection, it warns, “network overhead can reduce the latency gains”. The guide’s examples keep one socket open across turns and pass the previous response ID forward.
Plan for separate rate limits
GPT-6.1 Sol Ultrafast has “separate rate limits from Standard and Fast modes”, and OpenAI tells customers to check their organisation’s limits before increasing traffic. Standard GPT-6.1 Sol limits run from 5,000 requests and one million tokens per minute on the Build tier to 15,000 requests and 40 million tokens per minute on Grow. OpenAI has not published the Ultrafast equivalents, so test with real traffic before a launch.
Data residency for UK and European teams
GPT-6.1 Sol Ultrafast supports inference residency in the United States and in Europe, defined as the EEA plus Switzerland. GPT-6 Astra Ultrafast is US-only. The UK is not in the EEA, so British firms with strict UK-residency requirements should note that Europe here means EU processing, which carries the 10% regional uplift. For many UK businesses EU processing is acceptable, but record that decision in your data protection records.
Model details that do not change
The underlying model is the same. GPT-6.1 Sol has a 1,050,000-token context window, accepts up to 922,000 input tokens, writes up to 128,000 output tokens and has a knowledge cutoff of 30 April 2026. Its reasoning effort runs from low to max, with medium as the default. Speed tier and reasoning effort are separate settings, so change one at a time when you compare them.
Who Is Serving GPT-6.1 Sol Ultrafast?
Ultrafast began as a hardware story. When OpenAI first previewed the tier on 13 August for GPT-5.6 Sol, it ran on chips from Cerebras, the wafer-scale chipmaker with which OpenAI signed a 750-megawatt compute deal in January 2026. That preview promised up to 14 times Standard speed and about 750 tokens per second for select customers.
GPT-6.1 Sol Ultrafast may be different. Research firm SemiAnalysis said on X that it is “NOT running on Cerebras” and instead runs at a low batch size on Nvidia GPUs, as OfficeChai reported. OpenAI has not said which hardware serves the tier, and Cerebras did not comment.
Why low batch sizes cost more
A GPU serving many users at once keeps itself busy and makes each token cheap, but every user waits longer. Serving fewer users at a time does the opposite. OfficeChai noted that a 6x price premium fits the idea of reserving more expensive GPU capacity per request, while stressing that this is an inference. Cerebras shares hit a post-IPO low in early October, falling 20% in a week amid pressure from Nvidia and a lock-up expiry, CNBC reported.
How Ultrafast got here
| Date | Milestone |
|---|---|
| 30 July 2026 | Priority processing renamed Fast mode, up to 2.5x faster at 2x the price |
| 13 August 2026 | Ultrafast announced as a limited preview for GPT-5.6 Sol, up to 14x faster than Standard |
| 29 September 2026 | DevDay: GPT-6.1 Sol launched; Ultrafast added for GPT-6 Astra (US residency only) |
| 8 October 2026 | GPT-6.1 Sol Ultrafast rolls out in the API, Codex and ChatGPT Work, with US and EU residency |
| 14 October 2026 | GPT-5.5 retires from ChatGPT, ChatGPT Work and Codex (the API is unaffected) |
For the August preview and the Playground speed selector that followed, see our report on OpenAI’s Ultrafast API and speed selector.
When GPT-6.1 Sol Ultrafast Is Worth the Premium
Speed is only valuable when someone or something is waiting for it. The question for any team is not “is GPT-6.1 Sol Ultrafast faster?” but “what does each saved second earn us?”
Good fits
- Interactive coding with Codex, where an engineer waits for each change before making the next one.
- Customer-facing assistants where a slow reply loses the customer, and the volume is modest.
- Live demos, sales calls and support sessions where a pause is visible to a client.
- Agent loops with many short model calls between quick tool calls, run over WebSockets.
Poor fits
- Overnight batch jobs, document processing and evaluations, which belong on Batch or Flex at half the Standard price.
- Agent runs dominated by test suites, browsing or human review, where faster tokens barely move the total.
- Very high-volume, low-value requests, where a 6x token bill outweighs any gain in user satisfaction.
A simple test before you commit
Pick ten real tasks from your team’s week. Run each on Standard and on GPT-6.1 Sol Ultrafast with the same prompt and reasoning effort, and record time to a finished, accepted result, not just time to first output. Multiply the seconds saved by the cost of the person waiting, and compare it with the extra token spend. Our cloud cost optimisation team uses the same approach for any premium tier: measure the whole task, then decide.
What GPT-6.1 Sol Ultrafast Means for the Market
Speed is becoming a product line in its own right. OpenAI now sells the same GPT-6.1 Sol model at four price points, from half the Standard price on Batch to six times it on Ultrafast. Anthropic offers a premium fast mode for its Opus models, and Google sells its Flash models for speed and volume. Buyers increasingly choose a model and then choose how quickly they want it.
The near-term questions are capacity and hardware. If GPT-6.1 Sol Ultrafast becomes popular on GPUs, OpenAI must decide whether to keep reserving expensive capacity for it, raise the price, or move it to Cerebras chips. For a broader view of where each lab’s models sit, visit our AI models and tools hub.
Frequently Asked Questions About GPT-6.1 Sol Ultrafast
How much does GPT-6.1 Sol Ultrafast cost?
In the API, $12 per million input tokens and $60 per million output tokens for prompts up to 272,000 tokens, which is six times the Standard price. Cached input costs $0.60 and cache writes $15 per million tokens.
Can I use GPT-6.1 Sol Ultrafast on ChatGPT Plus?
No. In ChatGPT and Codex it is limited to Pro $500 and eligible Enterprise and Edu plans. OpenAI says other self-serve plans do not have access at launch, even with purchased credits.
Is GPT-6.1 Sol Ultrafast smarter than Standard?
No. It is the same model served faster. OpenAI has not published separate quality scores, and Kingy AI’s early test found no quality difference on its tasks.
Does GPT-6.1 Sol Ultrafast support EU data residency?
Yes. It supports inference residency in the United States and in Europe, meaning the EEA and Switzerland. Regional processing adds 10% to the price.
References
OpenAI Developers on X: Ultrafast is rolling out today for GPT-6.1 Sol
TestingCatalog on X: GPT-6.1 Sol Ultrafast is rolling out
OpenAI: ChatGPT and Codex changelog
OpenAI: GPT-6.1 Sol model page
OpenAI: Speed modes in Codex and ChatGPT Work
OpenAI: ChatGPT Work and Codex pricing
Kingy AI: GPT-6.1 Sol Ultrafast, is the time saved worth the price?
OfficeChai: GPT-6.1 Sol Ultrafast is not running on Cerebras, says SemiAnalysis
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.