4 trillion tokens pass through GMI Cloud’s inference platform every week, according to the Mountain View company, which disclosed the figure on 30 September 2026 as it announced $668 million in new financing. GMI Cloud rents out NVIDIA GPU capacity and runs AI models for other companies, several of which are themselves well-known AI platforms. The token count is the one number in the announcement that describes work actually being done today rather than money raised or revenue contracted.

A token is a small chunk of text, roughly three-quarters of an English word, and it is how most AI providers meter usage and charge for it. So a weekly token count is a rough measure of how much AI work a platform serves. But on its own, “4 trillion tokens” says little. Is that large? What could it be worth? Who is generating it?

This article puts the figure in context: what it means per second, who the customers are, how it compares with the biggest platforms, what it could be worth, and what the new money changes. Where we do arithmetic, we say so. For background on how token bills build up, see our AI token cost calculator guide and our explainer on FinOps for AI and GPU costs.

What 4 Trillion Tokens a Week Means

4 trillion tokens weekly gmi cloud new investment b fire engine spraying a water arc

The company’s press release, distributed through Business Wire, says only: “Its inference platform now processes approximately 4 trillion tokens per week.” It does not split input tokens (what customers send) from output tokens (what the models write), and it does not name the models.

Per day, per second, per year

Our arithmetic turns the weekly figure into other units. 4 trillion tokens a week is about 571 billion tokens a day. Spread evenly, that is about 6.6 million tokens every second. Over a year, at the same rate, it would be about 208 trillion tokens.

A factory’s worth of output

In November 2025, Reuters reported that GMI Cloud’s planned $500 million AI factory in Taiwan would house about 7,000 NVIDIA Blackwell GB300 GPUs in 96 racks, draw about 16 megawatts and be able to process “nearly 2 million tokens per second”. The company’s current average of 6.6 million per second is roughly 3.3 times that one site’s stated capacity, by our arithmetic. So the 4 trillion tokens must come from several sites, which fits GMI Cloud’s description of GPU infrastructure “across the United States and Asia-Pacific”.

UnitTokensHow we got it
Per weekAbout 4 trillionCompany figure
Per dayAbout 571 billionWeekly figure divided by 7
Per secondAbout 6.6 millionDaily figure divided by 86,400
Per yearAbout 208 trillionWeekly figure times 52
Taiwan factory capacityNearly 2 million per secondReuters, November 2025

Why the unit matters

Tokens are the common currency of AI services, which is why platforms quote them. They are also a flattering measure, because a single long document summary or a chatty coding agent can burn through tokens at a pace no human reader would. A rising count shows demand; it does not show profit.

What 4 Trillion Tokens Looks Like in Practice

4 trillion tokens weekly gmi cloud new investment c clawfoot bathtub overflowing with foam

Big token numbers are hard to picture, so we translated the weekly figure into everyday AI jobs. The job sizes below are our assumptions, chosen to be typical rather than precise.

Chats, documents and agent runs

A short customer-service chat with an AI assistant might use about 3,000 tokens, counting both the question and the reply. At that size, 4 trillion tokens would cover about 1.3 billion chats a week. Summarising a 40-page report might use about 30,000 tokens, which gives about 133 million summaries a week. A long coding-agent session that reads files, plans, edits and tests can easily use 2 million tokens, which gives about 2 million agent sessions a week.

Job (our assumed size)Tokens per jobJobs per week at 4 trillion tokens
Short support chat3,000About 1.3 billion
40-page report summary30,000About 133 million
Long coding-agent session2,000,000About 2 million

Why agents change the maths

The last row explains why token counts are climbing so fast across the industry. An agent does not answer once; it loops, rereading context and calling tools, and every loop costs tokens. The same 4 trillion tokens that could serve over a billion simple chats would serve only a couple of million heavy agent sessions. As more customers move from chatbots to agents, a platform’s token count can rise sharply even if its number of end users barely moves.

Who Is Behind the 4 Trillion Tokens

4 trillion tokens weekly gmi cloud new investment d zigzag marble run with rolling marbles

The customer list is the most revealing part of the release. It names Fireworks, Higgsfield, Nous Research, OpenRouter, Reflection, Cartesia, Trend Micro and Utopai Studios.

Platforms serving platforms

Several of those names sell AI to other developers. Fireworks runs a large inference platform of its own. OpenRouter routes developers’ requests across hundreds of models from many providers. Nous Research trains open models and agents, Reflection builds frontier open models, Cartesia makes voice models, Higgsfield and Utopai Studios work on AI video, and Trend Micro is a cybersecurity company that also invested in the round. In other words, a meaningful share of the 4 trillion tokens is probably wholesale work: GMI Cloud runs the hardware, and another brand sells the result to the end customer.

What a customer said

The release includes one customer quote. “Capacity that arrives late is capacity we can’t use,” said Chenyu Zhao, co-founder of Fireworks. “GMI Cloud has been one of our strongest and most reliable providers across NVIDIA GB200 and GB300 NVL72 systems.” That confirms Fireworks rents GMI Cloud’s rack-scale Blackwell systems, not just single GPUs.

A small mismatch in the coverage

The Korean outlet WOWTALE, citing founder and chief executive Alex Yeh’s post on X, said the announcement described customers without naming them, and matched “the lab behind the world’s most-used open-source AI agent” to OpenClaw. The Business Wire release does name customers, and the agent-lab name on its list is Nous Research, not OpenClaw. We use the release’s list.

Named customerType of company
FireworksInference and model customisation platform
OpenRouterRouter across hundreds of models
Nous ResearchOpen model and agent lab
ReflectionOpen frontier model lab
CartesiaVoice model company
HiggsfieldAI video platform
Utopai StudiosAI film studio
Trend MicroCybersecurity company; also an investor

How GPUs Turn Into 4 Trillion Tokens

4 trillion tokens weekly gmi cloud new investment e cornucopia spilling round tokens

An inference platform is a GPU cloud with software on top that keeps the chips busy. The economics depend on three things the release only hints at.

Rack-scale systems

Fireworks’ quote mentions NVIDIA GB200 and GB300 NVL72 systems. Each NVL72 rack links 72 Blackwell GPUs so that very large models can run across the whole rack as if it were one machine. That design suits serving big open models to many users at once, which is exactly what inference platforms sell.

Batching and utilisation

A GPU earns the most when it serves many requests together, a technique called batching. Idle chips still cost money for power, cooling and loan repayments, so utilisation is everything. In November 2025 Yeh told Reuters that GMI Cloud’s GPU utilisation was “almost full”. A full fleet is good for revenue but leaves little room for a sudden new customer, which is one reason the company keeps raising money for new capacity.

Hardware first, tokens second

GMI Cloud buys and installs the hardware, then either rents it out whole or runs models on it and meters the output. The 4 trillion tokens measure only the second kind of business. A customer that rents a full cluster and runs its own software may generate far more tokens than that, but they would not appear in GMI Cloud’s count.

4 Trillion Tokens Next to the Biggest Platforms

4 trillion tokens weekly gmi cloud new investment f jet engine on cradle stands

To judge the size of the number, we converted other public figures into tokens per week. Each comes from a different date and counts slightly different things, so treat the comparison as a sense of scale, not a league table.

The comparison figures

Google told investors on its second-quarter 2026 earnings call that its model APIs were “processing approximately 22 billion tokens per minute”, up from 16 billion a quarter earlier. That is about 222 trillion a week. Fireworks said in July 2026, announcing a $1.5 billion Series D at a $17.5 billion valuation, that it served “more than 40 trillion” tokens a day, about 280 trillion a week. OpenRouter said in May 2026, announcing a $113 million Series B led by CapitalG, that weekly volume had reached 25 trillion tokens.

Tokens per week, trillions (company figures converted by us; dates differ)

Fireworks, July 2026 (40T+ a day x 7): 280
Google model APIs, Q2 2026 (22B a minute): 222
OpenRouter, May 2026: 25
GMI Cloud, September 2026: 4

Small next to the leaders

On these numbers, 4 trillion tokens a week is about 1.4% of Fireworks’ weekly volume, 1.8% of Google’s API volume and 16% of OpenRouter’s May figure. GMI Cloud is not a top-tier token platform by volume, and it does not claim to be.

The double-counting problem

The comparison has a catch. Fireworks and OpenRouter are GMI Cloud customers. A token that a developer sends through OpenRouter to a model hosted by Fireworks on GMI Cloud hardware could appear in all three companies’ counts. Industry token totals therefore cannot be added up. For a buyer, the practical point is that the brand on your invoice may not own the GPUs doing the work.

What 4 Trillion Tokens Could Be Worth

The release gives two revenue numbers: contracted annual recurring revenue (ARR) of “more than $600 million”, more than nine times its level at the end of 2025, and live ARR in production that has grown “more than 4.5x” over the same period, with no dollar figure. It does not say how much of either comes from token-metered inference rather than GPU rental.

An illustration, not a disclosure

To show why that split matters, we priced the yearly run rate of about 208 trillion tokens at three blended rates per million tokens. These rates are our assumptions for illustration, covering the range from cheap open models to mid-priced ones; they are not GMI Cloud’s prices. At $0.10 per million, the 4 trillion tokens a week would be worth about $21 million a year. At $0.50, about $104 million. At $2.00, about $416 million.

Illustrative yearly value of 208 trillion tokens at assumed blended prices, US$ millions (our arithmetic)

Contracted ARR reported by the company: 600+
At $2.00 per million tokens: 416
At $0.50 per million tokens: 104
At $0.10 per million tokens: 21

What the illustration shows

Unless GMI Cloud charges well above typical open-model rates, metered tokens alone cannot account for $600 million of contracted revenue. The rest most likely comes from renting whole GPU clusters to customers like Fireworks, who then earn their own token revenue on top. That matches the company’s description of itself as “One Cloud for Compute, Inference, and Agents”, with GPU clusters listed first.

Contracted is not the same as live

Dividing $600 million by nine gives about $67 million, a rough guide to where contracted ARR stood at the end of 2025. The company did not publish that figure. Live ARR, the revenue actually running today, grew about half as fast in multiple terms, which is normal when contracts are signed before racks are installed.

The New Investment Behind the 4 Trillion Tokens

The token figure appeared in a funding announcement, and the money is aimed at making it bigger.

The round in brief

The $668 million comprises $223 million of Series B equity and a $445 million credit facility led by the Taiwanese bank CTBC. The equity round was led by ARCHIV, a new San Francisco firm focused on AI and robotics, with NVIDIA participating. Asia-Pacific investors include DSC Investment, Trend Micro, KB Investment, Kyobo Life and KT Corporation. No valuation was disclosed.

What it pays for

GMI Cloud says the money will expand capacity in the United States, Taiwan and the rest of Asia-Pacific, building on its Taiwan AI factory and a Japan sovereign AI initiative announced earlier this year. It will also fund “the continued growth of GMI Cloud’s inference services and strategic hiring”. In July 2026 the company had already committed $500 million of capital spending and announced a strategic collaboration with NVIDIA.

Why delivery dates matter

“In AI infrastructure, a delivery date is a promise,” Yeh said in the release. WOWTALE quoted him saying that “more than a quarter of expected data center capacity missed its completion date” last year, and that “every cluster we’ve committed has come online on schedule”. GMI Cloud credits its ties to Taiwan’s server supply chain for that record.

DateMilestone
2021Founded; headquartered in Mountain View, California
October 2024$82 million Series A
November 2025$500 million Taiwan AI factory announced, about 7,000 GB300 GPUs
July 2026$500 million capital spending commitment; NVIDIA collaboration
30 September 2026$668 million financing; about 4 trillion tokens a week disclosed

What Businesses Should Take From 4 Trillion Tokens

Most companies will never buy from GMI Cloud directly. But the announcement says something useful about the market they do buy from.

Your provider may be renting too

Inference platforms and model routers often run on GPU clouds such as GMI Cloud. That layering is not a problem in itself, but it affects where your data physically goes and whose outage becomes your outage. Ask a provider which infrastructure partners process your prompts and in which countries. Our coverage of Crusoe’s $3.9 billion raise describes the same layer from another angle.

Watch tokens per task, not per week

A platform’s weekly total tells you about its business. For your own costs, the number that matters is tokens per completed task. Agents and long-context tools can multiply it several times over, which is why cloud infrastructure planning for AI now starts with a usage forecast rather than a server count.

Regional capacity is a real option

GMI Cloud pitches capacity close to users in Asia-Pacific “under local data and compliance requirements”. For firms with customers in Taiwan, Japan or Korea, regional GPU clouds are a growing alternative to the largest US hyperscalers, and part of wider cloud computing choices about where data lives.

What We Still Don't Know About the 4 Trillion Tokens

The release is more precise than many funding announcements, but several gaps remain.

Which tokens

There is no split between input and output tokens, no list of models and no breakdown by customer. Output tokens usually cost several times more than input tokens, so the mix changes what the total is worth.

Which date

“Now processes” gives no week or period. The figure could be a recent peak or an average.

How much is inference revenue

Neither contracted nor live ARR is broken down between GPU rental and metered inference, and live ARR has no dollar figure.

How the IPO plan has moved

Reuters reported in November 2025 that GMI Cloud was looking at an initial public offering in two to three years, and that it planned a new 50-megawatt data centre in the United States. The September release mentions neither, so it is not clear whether either plan has changed.

FAQ: 4 Trillion Tokens and GMI Cloud

Which AI startup processes 4 trillion tokens a week?

GMI Cloud, a GPU cloud and inference provider founded in 2021 and based in Mountain View, California. It disclosed the figure on 30 September 2026.

Is 4 trillion tokens a week a lot?

It is about 6.6 million tokens a second. That is large for a young company, but small next to Google’s model APIs (about 222 trillion a week) or Fireworks (about 280 trillion a week), both of which published higher figures this year.

How much did GMI Cloud raise?

$668 million: $223 million of Series B equity led by ARCHIV with NVIDIA participating, plus a $445 million credit facility led by CTBC.

Who are GMI Cloud’s customers?

The release names Fireworks, Higgsfield, Nous Research, OpenRouter, Reflection, Cartesia, Trend Micro and Utopai Studios.

What is a token?

A small piece of text, roughly three-quarters of an English word, used by AI providers to measure and bill usage.

References and Further Reading