Snorkel AI raised $350 million in a Series E round announced on 22 September 2026, valuing the seven-year-old San Francisco company at $3.5 billion. That is close to three times the $1.3 billion it carried after its previous round in May 2025, and it arrives on the back of a revenue curve that is steeper than the valuation. The round was led by Insight Partners and S32.

The interesting number is not the valuation. It is what happened underneath it. When Snorkel AI last raised, its annualised revenue run-rate was in the low tens of millions. Chief executive Alex Ratner now says it has crossed $375 million, an eighteen-fold increase driven almost entirely by a data-as-a-service business the company only launched in September 2025. The company also expects to be profitable this year, which is not a sentence many AI companies can write in 2026.

This article covers what Snorkel AI actually sells, why the figures reported by different outlets disagree, how the valuation multiple has changed rather than simply grown, who else is competing for the same budgets, and what the whole episode says about where model capability is currently bottlenecked. The short version: the scarce input in frontier AI is no longer compute or talent. It is graded, adversarial, expert-level work that a model can be trained against.

What Snorkel AI Announced on 22 September 2026

snorkel ai valuation 3 5 billion training data boom d butter pat mould with one square recess

The announcement has four moving parts, and only one of them is the headline figure.

The round itself

Snorkel AI raised $350 million in Series E financing at a $3.5 billion post-money valuation. Insight Partners and S32 co-led. New participants named in the company’s own post include Third Point, March, Blumberg, Allegis, Standard VC and Frontline. Existing backers Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures and Factory all returned.

The revenue disclosure

Alongside the raise, Snorkel AI disclosed that its annualised revenue run-rate has crossed $375 million, representing eighteen-fold growth since it launched its data-as-a-service offering a year earlier. Reuters, working from its own sourcing, reported the figure as “crossed $350 million” from roughly $20 million a year before. Both framings describe the same trajectory.

The customer mix

The company sells to frontier AI labs, hyperscalers, large enterprises and the United States federal government. Coding data is its single biggest area of demand, which tracks with where the labs are spending: agentic coding is the benchmark suite every frontier release is now judged against.

The stated plan for the money

Ratner has said the capital goes towards hiring researchers and engineers, expanding enterprise and government operations, supporting third-party model evaluations, and pushing into new verticals and data modalities. Notably, evaluation-as-a-service for third parties is a different business from supplying training data to the labs, and it carries a different conflict profile.

The Snorkel AI Numbers, and Where They Disagree

snorkel ai valuation 3 5 billion training data boom e stepladder with three treads and a flat top platform

Three credible sources published three slightly different revenue figures on the same day. That is worth untangling rather than papering over.

The two run-rate figures

Snorkel AI’s own blog post and TechCrunch both give $375 million and eighteen-fold growth. Reuters gives $350 million, up from roughly $20 million a year earlier, which is a 17.5-fold increase. A company’s own post is the authoritative source for its own run-rate, so $375 million is the number to work from, with the Reuters figure best read as a slightly earlier or more conservative cut of the same data.

Why “gross” matters in this sector

TechCrunch adds a detail the other write-ups omit. Several of Snorkel AI’s fast-growing competitors report gross annualised revenue while paying 60 to 70 per cent of it straight out to the domain specialists doing the work. Snorkel AI books its human expert costs in cost of goods sold instead, so its headline figure and theirs are not directly comparable.

FigureSnorkel AI blogTechCrunchReuters
Raise$350m$350m$350m
Valuation$3.5bn$3.5bn$3.5bn
Run-rate$375m$375m$350m
Growth stated18x18-foldfrom ~$20m
Prior valuationnot stated$1.3bn, Series D$1.3bn, May 2025

The multiple went down, not up

Here is the arithmetic nobody put in a headline. In May 2025 Snorkel AI was worth $1.3 billion on a run-rate somewhere below $20 million, which is a revenue multiple on the order of 65 times. Today it is worth $3.5 billion on $375 million, which is 9.3 times. The valuation nearly tripled while the multiple fell by roughly 85 per cent.

Why that is the healthier shape

A company whose valuation grows more slowly than its revenue is being repriced towards its fundamentals, not away from them. That does not make $3.5 billion cheap, and a 9.3-times multiple on a services-heavy revenue line is still a growth price. But it is a materially different risk profile from the 2025 round, and it is the strongest single argument that this raise is demand-driven rather than sentiment-driven.

Annualised run-rate, scaled against the current $375m figure
Sept 2025, around $20m 5%
Reuters figure, $350m 93%
Company figure, $375m 100%
Each bar is the stated figure divided by $375m. The $20m bar rounds to 5%.

From Labelling Software to Data as a Service

snorkel ai valuation 3 5 billion training data boom f honey dipper standing upright in a squat pot

The business Snorkel AI is being valued on today is not the business it was founded to run.

The Stanford origin

Snorkel AI commercialised in 2019 after four years of research in the Stanford AI lab. The original product was weak supervision: instead of paying humans to label every example, you wrote labelling functions that programmatically generated noisy labels, then let the system de-noise and combine them. It was a clever answer to a 2019 problem.

Why software alone stopped being enough

Selling a platform means your customer still has to assemble the experts, design the tasks and run the quality process. Frontier labs did not want a platform for that. They wanted the finished artefact. So Snorkel AI moved up the stack from tooling to output, launching data as a service in September 2025 and booking eighteen-fold growth within twelve months.

What “finished dataset” actually means here

It means a delivered corpus or environment, quality-checked, targeted at a specific model weakness, with the human expertise already applied. The customer receives something they can train or evaluate against immediately. That is a services business with software margins where the automation works, and thin margins where it does not.

The expert network behind it

Reuters reports that Snorkel AI now draws on tens of thousands of specialists across fields including coding, law and medicine. That is a recruiting and vetting operation as much as a research one, and it is the part of the business that does not obviously compound.

What "Data 2.0" Means in Practice

snorkel ai valuation 3 5 billion training data boom b coffee grinder box with a round top hopper

Ratner’s funding post is framed around a distinction he calls Data 1.0 versus Data 2.0, and the distinction does real work.

Data 1.0 is a staffing problem

In the first phase, the constraint is volume. You need more labelled examples, and getting them is fundamentally about recruiting and coordinating people at scale. It is hard, but it is operationally hard rather than intellectually hard, and it scales with headcount and budget.

Data 2.0 is a research problem

In the second phase, volume is worthless. What matters is what Ratner calls “the right curriculum of extremely complex, precisely targeted, high quality data” — tasks pitched at the edge of what the model currently fails at. Producing that requires knowing where the model fails, which requires evaluation infrastructure, which is itself a research capability.

The tasks are deliberately beyond one human

Snorkel AI’s stated bar for a frontier training environment is a task that a senior engineer might struggle with over days or weeks. The environment has to encode nuanced goals, target the model’s actual error modes, resist sophisticated cheating, and survive extensive quality checks. Ratner’s post says this is beyond the ability of even the smartest human experts to develop alone at scale.

The compounding loop

The company describes building what it calls an RSI engine for data: specialised models accelerate human experts, human supervision improves those models, and the loop tightens. Whether that compounds in practice is the central bet in this valuation, because it is the only mechanism by which a data business escapes linear scaling with headcount.

The Competitive Field Snorkel AI Is Raising Into

snorkel ai valuation 3 5 billion training data boom c four nested mixing bowls stacked inside one another

This is a crowded sector with an unusual amount of capital moving through it.

Scale AI reset the comparables

In June 2025 Meta acquired a 49 per cent stake in Scale AI for $14.3 billion and hired its chief executive. That transaction did two things: it validated the category’s price, and it made Scale AI commercially awkward for every frontier lab competing with Meta. A material share of Snorkel AI’s opportunity is downstream of that awkwardness.

The fast followers

Mercor reports roughly $2 billion in gross annualised revenue. Handshake passed $1 billion in 2026. Micro1 reports a $500 million gross run-rate. Surge AI remains a significant independent. On gross revenue alone, Snorkel AI is not the largest player in its own category.

CompanyStated revenueBasisNotable
Snorkel AI$375m run-rateExpert cost in COGSExpects profit this year
Mercor~$2bnGross annualised60-70% paid to experts
Handshake$1bn in 2026GrossPivoted from campus hiring
Micro1$500m run-rateGrossFastest recent riser
Scale AINot disclosedPrivateMeta bought 49% for $14.3bn

What Snorkel AI is arguing differentiates it

The pitch is margin structure and research depth rather than scale. If most of the sector is a marketplace that routes work to contractors, and Snorkel AI is a research shop that uses automation to make each expert hour go further, then the same revenue line means something different on each balance sheet. Profitability this year, if it lands, is the evidence for that claim.

The obvious counter-argument

Automation advantages in data work have historically been competed away quickly, because the techniques publish. Snorkel AI’s moat has to be the accumulated evaluation infrastructure and the expert network, not any single method.

Why Reinforcement Learning Environments Are the Expensive Part

The phrase doing most of the work in this funding round is “RL environments”, and it deserves unpacking.

An environment is not a dataset

A dataset is a fixed set of examples with known answers. An environment is an executable world with a task, a state, and a grader that can score an attempt. Training a model with reinforcement learning against an environment means the model can try, fail, and be scored thousands of times without a human in the loop for each attempt.

Building the grader is the hard bit

Anyone can write a hard task. Writing a grader that recognises a correct solution it has never seen, rejects a plausible wrong one, and cannot be gamed by a model that has learned to satisfy the letter of the check is genuinely difficult. Snorkel AI’s own research blog has published on exactly this failure surface.

The cheating problem is real

Ratner’s post explicitly lists resistance to “sophisticated cheating attempts” as a design requirement. Models optimised hard against a grader will find the grader’s blind spots, and an environment that can be gamed actively teaches the wrong behaviour rather than merely failing to teach the right one.

Why coding data leads demand

Code is the domain where automated grading is most tractable: tests either pass or they do not. That makes coding environments the cheapest high-quality signal available, which is why Snorkel AI names coding as its largest demand area and why every frontier lab’s headline benchmarks are now agentic coding suites.

What the Revenue Number Does and Does Not Tell You

Eighteen-fold growth in twelve months is a real signal, but it is a signal about a specific thing.

It measures lab spending, not enterprise adoption

The buyers driving this curve are frontier labs and hyperscalers with enormous training budgets and a small number of very large contracts. That is a concentrated revenue base by construction, and concentration cuts both ways.

Lab budgets are correlated

If frontier training spend slows — because of a capability plateau, a funding squeeze, or an industry agreement to pace development — it slows across most of Snorkel AI’s largest customers simultaneously. The federal government and enterprise lines are the diversification, which is presumably why the funding post names them.

Revenue multiple at each round, scaled against the 2025 figure
May 2025: $1.3bn on under $20m, about 65x 100%
Sept 2026: $3.5bn on $375m, 9.3x 14%
9.3 divided by 65 is 0.14, so the second bar is 14% of the first. The multiple compressed by about 86%.

The profitability claim is the one to watch

Reuters reports that Snorkel AI expects profitability this year. In a sector where the largest players run marketplace economics and pay most of their revenue out, a profitable data supplier would be a genuinely different animal. It is also the claim most easily checked in twelve months.

What Snorkel AI's Round Means If You Buy AI Rather Than Build It

Almost no one reading this is training a frontier model. The round still says something useful.

Capability is bought, not just computed

The prevailing story of the last three years was that capability came from scale. This round is evidence that the marginal capability gain now comes from curriculum: which specific hard tasks the model was trained against. That is why two models on similar compute budgets can differ sharply on the work you actually care about.

Your evaluation data is worth more than you think

If graded, domain-specific, adversarial task sets are the scarce input at the frontier, they are also the scarce input inside your organisation. Most businesses deploying autonomous AI agents have no equivalent of a grader for the work those agents do, which is exactly why reliability is hard to argue about internally.

Benchmarks now have a supply chain

When a lab reports a score, some of the data behind that capability was bought from a small number of suppliers who also sell to the lab’s competitors. That is not a scandal, but it does mean cross-lab benchmark gaps can narrow for commercial reasons rather than research ones. Keeping track of which AI models and tools lead on which task is now partly a procurement question.

The practical takeaway

Judge a model on your own work, with your own graded examples, rather than on published suites. That advice has always been sensible; a $3.5 billion valuation for the company that builds those graded examples is a reasonable indication it is now necessary.

How Snorkel AI Produces Data at This Scale

The operational question a $375 million run-rate raises is how the work actually gets done, because pure human labour would not get you there.

The hybrid production model

Snorkel AI describes a hybrid approach that pairs synthetic and automated generation with subject matter experts. The machine drafts and varies; the expert judges, corrects and signs off. Neither half produces frontier-grade data alone, and Ratner has been explicit that both are required.

The quote that defines the strategy

Asked about the balance, Ratner told Reuters: “Our strong view is that 100% of the data that labs will get value out of will have some human input in the foreseeable future. But 100% of that data will have to use synthetic and automated approaches to keep up with this complexity.” Both halves of that sentence are load-bearing.

Why the expert cannot be removed

Oversight, alignment and safety work all require a human to define what correct looks like. A fully automated pipeline can only encode a standard someone already wrote down, which is precisely what does not exist for tasks at the edge of model capability.

Why the automation cannot be removed either

Complexity is rising faster than expert hours can be bought. Snorkel AI’s argument is that automation multiplies each expert hour rather than replacing it, and the specialist network of tens of thousands across coding, law and medicine is the input that multiplication acts on.

What the investors bought

Andy Harrison, a partner at S32, framed the thesis bluntly: “Data is becoming more rare, more specialized, more difficult to find. If you want to train the most frontier, complex and capable models, now you need superior data.” That is a scarcity argument, and scarcity arguments price well.

What Could Go Wrong With the Snorkel AI Thesis

Three risks sit underneath the round, and none of them is exotic.

Customer concentration

A revenue base built on frontier labs is a handful of very large contracts. Losing or shrinking one is a material event, and Snorkel AI does not disclose how concentrated it is.

The labs building it in house

Every frontier lab has the capability to build its own data pipelines, and several already do for their most sensitive domains. Snorkel AI’s defence is speed and specialist supply rather than anything a lab could not replicate given time.

Methods publish

Data techniques diffuse quickly through the research literature — Snorkel AI’s own original weak-supervision work is the obvious example. Any automation advantage has a half-life, which is why the durable asset has to be the expert network and the accumulated evaluation infrastructure rather than the method.

Frequently Asked Questions About Snorkel AI

What does Snorkel AI actually sell?

Finished training and evaluation datasets, plus reinforcement learning environments, delivered to customers rather than produced by them. It began as weak-supervision labelling software and moved to data as a service in September 2025.

Who led the Series E?

Insight Partners and S32 co-led the $350 million round. Existing investors including Addition, Lightspeed, Greylock, GV and Wells Fargo participated, alongside new names such as Third Point and Blumberg.

Is Snorkel AI profitable?

Not yet, but Reuters reports the company expects to reach profitability this year. Its accounting differs from several competitors in that expert costs sit in cost of goods sold rather than being netted against gross revenue.

How does Snorkel AI compare with Scale AI?

Scale AI is larger and older, and Meta acquired a 49 per cent stake in it for $14.3 billion in June 2025. That stake makes Scale AI a complicated supplier for labs competing with Meta, which has helped independents including Snorkel AI.

Why is coding data in such demand?

Because code can be graded automatically by running tests, making it the cheapest domain in which to build reliable training environments. It is also the capability frontier labs currently compete on hardest.

References