Pangram has become the name publishers, professors and platforms reach for when they want to know whether a piece of writing came from a human or a machine. The 24-person startup — headquartered, improbably, above a Popeyes in Brooklyn — has been called the gold standard of AI detection, and its verdicts have already cancelled a book deal, embarrassed The New York Times and cast doubt on a major literary prize. WIRED’s Lexi Pandell profiled the company on 2 September 2026 and asked the question everyone in publishing is now asking: how much can it be trusted?

The stakes are not academic. A Pangram score is a number, and numbers travel: a screenshot of “97 percent AI-generated” can end a career before the accused writer has even seen the report. Yet independent benchmarks really do rank Pangram far ahead of its rivals, with a false-positive rate competitors cannot match. If you have ever tried to spot AI writing by eye, you will know why institutions want an automated answer.

This article examines both sides of the ledger: what Pangram is, how it works, what the independent evidence shows, where it has gone wrong, and — most practically — when you should and should not rely on its verdicts. The original reporting is in WIRED’s investigation, which we draw on throughout alongside peer benchmarks and the company’s own technical claims.

What Is Pangram? The 24-Person Startup Policing the Written Word

pangram ai detection gold standard trust b solid meter box round blank dial

Pangram is an AI-detection startup that promises to estimate how much AI was involved in generating any given text, returning a best-guess percentage rather than a binary yes or no. It has raised $13 million to date — roughly 0.0072 percent of what OpenAI has raised — and was virtually unknown until this year. Today it caters to education, the legal field and recruitment, though creative writing makes up the largest segment of its training text.

From Checkfor.ai to publishing’s AI police

Cofounders Max Spero and Bradley Emi met as undergraduates at Stanford. Spero went on to Google, where he worked on FLoC, the ill-fated cookie-replacement technology, and later the autonomous-vehicle company Nuro. Emi passed through Tesla’s Autopilot team and the AI biotech firm Absci. After ChatGPT launched in 2022, the pair saw the coming flood of machine-generated text as a business opportunity, founding Checkfor.ai in 2023 and renaming it Pangram a year later.

Funding, growth and the Substack deal

The company’s rise has been rapid. It closed a $4 million seed round in June 2025, then raised $9 million in July 2026 in a round led by Menlo Ventures, announced alongside its newest model, Pangram 4. In late July, Substack said it would integrate the detector into its platform so readers could check posts for possible AI use. The timeline below shows how quickly a tiny startup became publishing’s de facto referee.

DateMilestoneWhy it mattered
2023Founded as Checkfor.aiEntered a field already crowded by Originality.ai, GPTZero and Turnitin
2024Renamed PangramEarly independent tests marked it as a front-runner
June 2025$4M seed round closedFunded the push into education and enterprise
August 2025NBER working paper publishedIndependent benchmark named it the clear accuracy leader
January 2026Shy Girl controversyA 78% score preceded Hachette cancelling the novel’s release
July 2026$9M raise, Pangram 4 launch, Substack integrationDetection moved from back office to reader-facing feature

How Pangram AI Detection Works

pangram ai detection gold standard trust c solid dome bell round top knob

Pangram’s approach differs from the first wave of detectors, which leaned on crude statistical signals and earned a reputation for wrongly accusing honest writers. Early tools — including the classifier OpenAI itself retired in 2023 after it caught only 26 percent of AI text while mislabelling human writing 9 percent of the time — made educators rightly sceptical. Even mainstream writing tools like Grammarly complicated the picture, since AI-assisted editing blurs the line detectors try to draw.

Synthetic mirroring

The core technique is what the company calls synthetic mirroring. Pangram takes a piece of human writing and has today’s leading AI models generate a close match, teaching its classifier the difference between the two on near-identical content. Because the model is trained on natural language processing signals rather than surface clichés, it is harder to fool with light paraphrasing than its predecessors.

Hard negative mining

The second technique is hard negative mining: the company actively hunts its own false positives — human texts its model finds hardest to classify — then mirrors and reintegrates them into training. Spero says all its datasets are properly licensed, “which I think is kind of rare in the AI world today.” The result, he argues, is a detector that keeps improving precisely where it previously failed.

A percentage, not a verdict

Crucially, the tool returns a probability-flavoured percentage, and the company says it deliberately errs on the side of calling borderline text human-written. “This is an intentional choice that we made,” Spero told WIRED. That choice keeps false accusations rare — and, as we will see, it also means a meaningful share of AI-assisted work slips through undetected.

The Accuracy Record: What Independent Testing Shows

pangram ai detection gold standard trust d three identical upright cylinders

Pangram’s reputation does not rest on marketing alone. The strongest evidence comes from an August 2025 NBER working paper, Artificial Writing and Automated Detection, by Brian Jabarian and Alex Imas of the University of Chicago’s Booth School. The researchers tested four detectors against a corpus of 1,992 verified pre-2020 human texts and 1,992 matching AI-generated texts spanning genres, lengths and models.

The NBER benchmark results

The findings were unusually one-sided. Pangram achieved near-zero false-positive and false-negative rates, robust across generating models, threshold rules, ultra-short passages and even so-called humanizer tools such as StealthGPT. It was the only detector to meet a stringent policy cap — a false-positive rate at or below 0.5 percent — without sacrificing its ability to catch AI text. On long and medium passages its false-positive rate was essentially zero; on short passages it ticked up but never above 1 percent.

DetectorFalse positives (medium/long text)False positives (short text)Meets 0.5% policy cap
PangramEssentially zeroUnder 1%Yes — the only one
GPTZeroAround 1% or lessUp to ~2.4%No
Originality.aiAround 1% or lessUp to ~3%No
RoBERTa (open source)Very highUp to ~69% — flags most human textNo

The gap on short passages is worth seeing side by side, because short excerpts — a paragraph pasted from a novel, a suspect student answer — are exactly what real accusers scan. The bars below chart the worst-case share of human-written short texts each detector wrongly flagged in the Booth testing.

Worst-case false flags on short human texts (NBER benchmark)
RoBERTa (open source) 69%
Originality.ai 3%
GPTZero 2.4%
Pangram 1%

The non-native English question

A famous 2023 Stanford study by Liang and colleagues found that early GPT detectors flagged an average of 61.3 percent of essays by non-native English speakers as AI-generated — a devastating bias. Pangram later ran the same TOEFL dataset through its own model and reported a 0.00 percent false-positive rate, though that retest is the company’s own figure rather than an independent one. The distinction matters, because bias claims are central to the case against detection, as we cover below.

The Scandals: When a Pangram Score Ends a Book Deal

pangram ai detection gold standard trust e solid upright funnel wide mouth

Pangram’s public profile was built less on benchmarks than on a string of literary scandals — each one amplified by the company itself.

Shy Girl and the Hachette cancellation

In January, Reddit and YouTube speculation suggested the author Mia Ballard may have used AI to write Shy Girl, a self-published novel Hachette had picked up for traditional publication. Ballard denied it. Spero ran the manuscript through his own tool and posted on X that the book was 78 percent AI-generated; Hachette later cancelled the release. WIRED’s reporting adds two uncomfortable wrinkles: the score reached The New York Times through a chain involving a Pangram account executive, and critics noted Spero’s copy of the manuscript came from a pirating website. His response: “I hadn’t looked too closely. I just put it straight into Pangram.”

The Times, the prize and a $2.4 million thriller

A barrage of accusations followed, summarised in the table below. The New York Times was called out for an apparently AI-generated Modern Love instalment. A Commonwealth Short Story Prize winner scored 100 percent — after which the company scanned every winner since 2012 and flagged three more. The novel Daggermouth scored 60 percent, and Call Me, I’ll Hide the Body, a thriller that had sold for $2.4 million, scored 97 percent.

WorkPangram scoreOutcome
Shy Girl (Mia Ballard)78%Hachette cancelled the release
NYT Modern Love column100%Public embarrassment for the paper
Commonwealth Short Story Prize winner100%Prize scrutiny; three past winners also flagged
Daggermouth60%Public accusation via academic dataset
Call Me, I’ll Hide the Body97%$2.4M deal engulfed in controversy

The conflict-of-interest web

WIRED’s profile also maps a cosy network around the company. Tuhin Chakrabarty, the Stony Brook professor whose scan of 14,419 self-published novels found nearly 20 percent had substantial AI-detection scores, receives free API credits from the company and calls Spero a close friend. Chakrabarty’s partner, Todd Shuster, co-CEO of the literary agency Aevitas, both uses the detector on manuscripts and consults for the company, making introductions to publishers. None of this invalidates the scores — but the industry’s referee is unusually entangled with the players.

The Case for Trusting Pangram

pangram ai detection gold standard trust f tall stack blank paper sheets

Sceptics of AI detection often cite the failures of 2023-era tools. The honest reading of the current evidence is that this generation is different, and one product in particular.

Independent numbers, not vendor claims

The NBER benchmark is the key exhibit: near-zero error rates, robust to humanizers like StealthGPT and to passages as short as 50 words, from researchers with no stake in the outcome. No competitor cleared the same bar. When GPT-5 launched, Spero claimed his was the only detector that reliably identified its output without explicit retraining. The company now claims a false-positive rate of 0.0041 percent with Pangram 4 — about 1 in 24,000 — down from a previously reported 0.01 percent.

A culture of showing its work

Unlike most rivals, the company publishes technical reports on its models and training methods, including a peer-reviewed technical paper on arXiv. Spero has offered cash bounties to writers who can prove they were wrongly flagged; so far, he says, nobody has taken him up on it. “Everybody else is sort of this black box,” he told WIRED. “Our strategy here is really just build trust through transparency.”

Industry adoption is accelerating

Substack’s integration puts detection in front of readers, not just gatekeepers. A Penguin Random House spokesperson confirmed its editors may use approved detection tools as “one additional means” of identifying AI content. Agents quietly screen submissions. And the demand is real: a Gotham Ghostwriters poll of 1,481 working writers found 61 percent use AI tools and 7 percent have published AI-generated text. Someone was always going to referee this; the evidence says this referee is the most accurate available.

The Case for Caution

The counter-case does not require the detector to be bad at statistics. It requires only that rare errors, multiplied by enormous scale, land on real people.

The base-rate problem

A 0.0041 percent false-positive rate sounds negligible — until you multiply. At that claimed rate, one million scans of genuinely human writing would produce roughly 41 false accusations; at the older 0.01 percent rate, about 100. Turnitin’s self-reported document-level rate of “under 1 percent” implies up to 10,000 per million. Substack-scale deployment means millions of scans, and each false flag is a human being facing an unfalsifiable charge.

False accusations per one million human texts scanned (claimed rates)
Turnitin (under 1% claim) up to 10,000
Pangram 3.x (0.01%) 100
Pangram 4 (0.0041%) 41

Humanizers punch a hole in the wall

A Notre Dame working paper, “Why AI Detection Fails for Academic Integrity,” found that after AI text was passed through a humanizer tool, the 3.2 model caught it less than 4 percent of the time — while flagging lightly AI-edited human abstracts 64 to 80 percent of the time. That is close to the worst possible combination for academic integrity. An entire industry of rewriting tools — we have reviewed Go Humanize AI and Rehumanize.io — exists to exploit exactly this gap, and The Atlantic reported similar results testing humanized text earlier this year.

Short texts, shifting verdicts

The company itself discloses that accuracy degrades under 100 words, that the median input is only about 350 words, and that an identical passage can receive different verdicts scanned alone versus in context. A pasted excerpt from a novel is precisely the short, decontextualised input where the tool is weakest — and precisely how public accusations happen.

Bias, and who gets accused

“I don’t think AI detectors work,” Sam Illingworth, a professor of critical AI literacy at Edinburgh Napier University, told WIRED. “I think that detectors are prejudiced against certain people.” The company disputes this with internal research, but the optics are stark: all three of publishing’s biggest detection scandals — Shy Girl, Daggermouth and Call Me — involved writers of color. As Regina Brooks, president of the Association of American Literary Agents, notes, whose work gets scanned and publicly accused is a human choice, and human choices carry bias.

The blind spot nobody screenshots

By the company’s own measure, human essays heavily rewritten by AI are still classified as human-written 41.37 percent of the time — the price of erring toward “human” on borderline text. The gold standard, in other words, systematically under-detects the most common real-world use of AI: heavy assistance rather than wholesale generation.

Should You Trust Pangram? A Practical Answer

The evidence supports a narrow conclusion: Pangram is the most accurate AI detector independently tested, and it is still not an oracle. Treat the score as strong probabilistic evidence, never as proof.

ScenarioHow much weight to give the scoreWhy
Long document, very high scoreHigh — corroborate, then actFalse positives on long text were essentially zero in independent testing
Short excerpt, borderline scoreLowAccuracy drops under 100 words; verdicts shift with context
A clean “human” verdictLow to moderateHumanizer tools evade detection over 96% of the time
Grounds for public accusationNever on the score aloneEven a 1-in-24,000 error rate produces real victims at scale

If you are a writer

Keep receipts. Drafts, version history, notes and research trails are the evidence that rebuts a false flag — the accusation is statistical, but your paper trail is not. Where platforms disclose scores, assume your work will be scanned and consider disclosing AI assistance up front; Spero’s own framing is that deception, not use, is what ends careers.

If you run a school, publisher or business

Write a policy before you buy a scanner. Decide what score triggers what process, who reviews it, and what corroboration is required — a detection score should open a conversation, not close one, echoing Penguin Random House’s “not determinative” stance. Detection also pairs with provenance: watermarking systems like Google’s SynthID and the EU’s incoming AI content-label rules attack the same problem from the supply side. Our AI governance framework for SMEs covers how to fold both into a workable policy.

FAQ

Is Pangram actually the most accurate AI detector?

On the independent evidence, yes. The NBER benchmark found it was the only detector among four tested to combine near-zero false positives with near-zero false negatives, and the only one meeting a 0.5 percent false-positive policy cap. That is a genuine, measured lead — not marketing.

Can Pangram be fooled?

Yes. Humanizer tools that rewrite AI output evaded the 3.2 model more than 96 percent of the time in Notre Dame’s testing, and writers have deliberately mimicked AI style to trigger false flags. Detection is an arms race, and evasion currently has the cheaper weapons.

What does a Pangram percentage actually mean?

It is a model’s best-guess estimate of how much AI was involved in generating the text — not a measured fact. The tool deliberately errs toward calling borderline text human, which keeps false accusations rare but lets heavily AI-assisted writing pass as human 41.37 percent of the time by the company’s own figures.

Should schools act on a score alone?

No. The company’s disclosed weaknesses — short texts, context-dependent verdicts, humanizer evasion — plus the bias concerns raised by researchers mean a score should trigger review and conversation, never automatic penalties. That is also how Penguin Random House says it treats detection internally.

Who is behind Pangram?

Stanford graduates Max Spero (ex-Google, Nuro) and Bradley Emi (ex-Tesla Autopilot, Absci) founded the company as Checkfor.ai in 2023. It has 24 employees, $13 million in funding led most recently by Menlo Ventures, and counts Substack as its highest-profile integration.

References