Pangram has become the name publishers, professors and platforms reach for when they want to know whether a piece of writing came from a human or a machine. The 24-person startup — headquartered, improbably, above a Popeyes in Brooklyn — has been called the gold standard of AI detection, and its verdicts have already cancelled a book deal, embarrassed The New York Times and cast doubt on a major literary prize. WIRED’s Lexi Pandell profiled the company on 2 September 2026 and asked the question everyone in publishing is now asking: how much can it be trusted?
The stakes are not academic. A Pangram score is a number, and numbers travel: a screenshot of “97 percent AI-generated” can end a career before the accused writer has even seen the report. Yet independent benchmarks really do rank Pangram far ahead of its rivals, with a false-positive rate competitors cannot match. If you have ever tried to spot AI writing by eye, you will know why institutions want an automated answer.
This article examines both sides of the ledger: what Pangram is, how it works, what the independent evidence shows, where it has gone wrong, and — most practically — when you should and should not rely on its verdicts. The original reporting is in WIRED’s investigation, which we draw on throughout alongside peer benchmarks and the company’s own technical claims.
Table of contents
- What Is Pangram? The 24-Person Startup Policing the Written Word
- How Pangram AI Detection Works
- The Accuracy Record: What Independent Testing Shows
- The Scandals: When a Pangram Score Ends a Book Deal
- The Case for Trusting Pangram
- The Case for Caution
- Should You Trust Pangram? A Practical Answer
- FAQ
- References
What Is Pangram? The 24-Person Startup Policing the Written Word
Pangram is an AI-detection startup that promises to estimate how much AI was involved in generating any given text, returning a best-guess percentage rather than a binary yes or no. It has raised $13 million to date — roughly 0.0072 percent of what OpenAI has raised — and was virtually unknown until this year. Today it caters to education, the legal field and recruitment, though creative writing makes up the largest segment of its training text.
From Checkfor.ai to publishing’s AI police
Cofounders Max Spero and Bradley Emi met as undergraduates at Stanford. Spero went on to Google, where he worked on FLoC, the ill-fated cookie-replacement technology, and later the autonomous-vehicle company Nuro. Emi passed through Tesla’s Autopilot team and the AI biotech firm Absci. After ChatGPT launched in 2022, the pair saw the coming flood of machine-generated text as a business opportunity, founding Checkfor.ai in 2023 and renaming it Pangram a year later.
Funding, growth and the Substack deal
The company’s rise has been rapid. It closed a $4 million seed round in June 2025, then raised $9 million in July 2026 in a round led by Menlo Ventures, announced alongside its newest model, Pangram 4. In late July, Substack said it would integrate the detector into its platform so readers could check posts for possible AI use. The timeline below shows how quickly a tiny startup became publishing’s de facto referee.
| Date | Milestone | Why it mattered |
|---|---|---|
| 2023 | Founded as Checkfor.ai | Entered a field already crowded by Originality.ai, GPTZero and Turnitin |
| 2024 | Renamed Pangram | Early independent tests marked it as a front-runner |
| June 2025 | $4M seed round closed | Funded the push into education and enterprise |
| August 2025 | NBER working paper published | Independent benchmark named it the clear accuracy leader |
| January 2026 | Shy Girl controversy | A 78% score preceded Hachette cancelling the novel’s release |
| July 2026 | $9M raise, Pangram 4 launch, Substack integration | Detection moved from back office to reader-facing feature |
How Pangram AI Detection Works
Pangram’s approach differs from the first wave of detectors, which leaned on crude statistical signals and earned a reputation for wrongly accusing honest writers. Early tools — including the classifier OpenAI itself retired in 2023 after it caught only 26 percent of AI text while mislabelling human writing 9 percent of the time — made educators rightly sceptical. Even mainstream writing tools like Grammarly complicated the picture, since AI-assisted editing blurs the line detectors try to draw.
Synthetic mirroring
The core technique is what the company calls synthetic mirroring. Pangram takes a piece of human writing and has today’s leading AI models generate a close match, teaching its classifier the difference between the two on near-identical content. Because the model is trained on natural language processing signals rather than surface clichés, it is harder to fool with light paraphrasing than its predecessors.
Hard negative mining
The second technique is hard negative mining: the company actively hunts its own false positives — human texts its model finds hardest to classify — then mirrors and reintegrates them into training. Spero says all its datasets are properly licensed, “which I think is kind of rare in the AI world today.” The result, he argues, is a detector that keeps improving precisely where it previously failed.
A percentage, not a verdict
Crucially, the tool returns a probability-flavoured percentage, and the company says it deliberately errs on the side of calling borderline text human-written. “This is an intentional choice that we made,” Spero told WIRED. That choice keeps false accusations rare — and, as we will see, it also means a meaningful share of AI-assisted work slips through undetected.
The Accuracy Record: What Independent Testing Shows
Pangram’s reputation does not rest on marketing alone. The strongest evidence comes from an August 2025 NBER working paper, Artificial Writing and Automated Detection, by Brian Jabarian and Alex Imas of the University of Chicago’s Booth School. The researchers tested four detectors against a corpus of 1,992 verified pre-2020 human texts and 1,992 matching AI-generated texts spanning genres, lengths and models.
The NBER benchmark results
The findings were unusually one-sided. Pangram achieved near-zero false-positive and false-negative rates, robust across generating models, threshold rules, ultra-short passages and even so-called humanizer tools such as StealthGPT. It was the only detector to meet a stringent policy cap — a false-positive rate at or below 0.5 percent — without sacrificing its ability to catch AI text. On long and medium passages its false-positive rate was essentially zero; on short passages it ticked up but never above 1 percent.
| Detector | False positives (medium/long text) | False positives (short text) | Meets 0.5% policy cap |
|---|---|---|---|
| Pangram | Essentially zero | Under 1% | Yes — the only one |
| GPTZero | Around 1% or less | Up to ~2.4% | No |
| Originality.ai | Around 1% or less | Up to ~3% | No |
| RoBERTa (open source) | Very high | Up to ~69% — flags most human text | No |
The gap on short passages is worth seeing side by side, because short excerpts — a paragraph pasted from a novel, a suspect student answer — are exactly what real accusers scan. The bars below chart the worst-case share of human-written short texts each detector wrongly flagged in the Booth testing.
The non-native English question
A famous 2023 Stanford study by Liang and colleagues found that early GPT detectors flagged an average of 61.3 percent of essays by non-native English speakers as AI-generated — a devastating bias. Pangram later ran the same TOEFL dataset through its own model and reported a 0.00 percent false-positive rate, though that retest is the company’s own figure rather than an independent one. The distinction matters, because bias claims are central to the case against detection, as we cover below.
The Scandals: When a Pangram Score Ends a Book Deal
Pangram’s public profile was built less on benchmarks than on a string of literary scandals — each one amplified by the company itself.
Shy Girl and the Hachette cancellation
In January, Reddit and YouTube speculation suggested the author Mia Ballard may have used AI to write Shy Girl, a self-published novel Hachette had picked up for traditional publication. Ballard denied it. Spero ran the manuscript through his own tool and posted on X that the book was 78 percent AI-generated; Hachette later cancelled the release. WIRED’s reporting adds two uncomfortable wrinkles: the score reached The New York Times through a chain involving a Pangram account executive, and critics noted Spero’s copy of the manuscript came from a pirating website. His response: “I hadn’t looked too closely. I just put it straight into Pangram.”
The Times, the prize and a $2.4 million thriller
A barrage of accusations followed, summarised in the table below. The New York Times was called out for an apparently AI-generated Modern Love instalment. A Commonwealth Short Story Prize winner scored 100 percent — after which the company scanned every winner since 2012 and flagged three more. The novel Daggermouth scored 60 percent, and Call Me, I’ll Hide the Body, a thriller that had sold for $2.4 million, scored 97 percent.
| Work | Pangram score | Outcome |
|---|---|---|
| Shy Girl (Mia Ballard) | 78% | Hachette cancelled the release |
| NYT Modern Love column | 100% | Public embarrassment for the paper |
| Commonwealth Short Story Prize winner | 100% | Prize scrutiny; three past winners also flagged |
| Daggermouth | 60% | Public accusation via academic dataset |
| Call Me, I’ll Hide the Body | 97% | $2.4M deal engulfed in controversy |
The conflict-of-interest web
WIRED’s profile also maps a cosy network around the company. Tuhin Chakrabarty, the Stony Brook professor whose scan of 14,419 self-published novels found nearly 20 percent had substantial AI-detection scores, receives free API credits from the company and calls Spero a close friend. Chakrabarty’s partner, Todd Shuster, co-CEO of the literary agency Aevitas, both uses the detector on manuscripts and consults for the company, making introductions to publishers. None of this invalidates the scores — but the industry’s referee is unusually entangled with the players.
The Case for Trusting Pangram
Sceptics of AI detection often cite the failures of 2023-era tools. The honest reading of the current evidence is that this generation is different, and one product in particular.
Independent numbers, not vendor claims
The NBER benchmark is the key exhibit: near-zero error rates, robust to humanizers like StealthGPT and to passages as short as 50 words, from researchers with no stake in the outcome. No competitor cleared the same bar. When GPT-5 launched, Spero claimed his was the only detector that reliably identified its output without explicit retraining. The company now claims a false-positive rate of 0.0041 percent with Pangram 4 — about 1 in 24,000 — down from a previously reported 0.01 percent.
A culture of showing its work
Unlike most rivals, the company publishes technical reports on its models and training methods, including a peer-reviewed technical paper on arXiv. Spero has offered cash bounties to writers who can prove they were wrongly flagged; so far, he says, nobody has taken him up on it. “Everybody else is sort of this black box,” he told WIRED. “Our strategy here is really just build trust through transparency.”
Industry adoption is accelerating
Substack’s integration puts detection in front of readers, not just gatekeepers. A Penguin Random House spokesperson confirmed its editors may use approved detection tools as “one additional means” of identifying AI content. Agents quietly screen submissions. And the demand is real: a Gotham Ghostwriters poll of 1,481 working writers found 61 percent use AI tools and 7 percent have published AI-generated text. Someone was always going to referee this; the evidence says this referee is the most accurate available.
The Case for Caution
The counter-case does not require the detector to be bad at statistics. It requires only that rare errors, multiplied by enormous scale, land on real people.
The base-rate problem
A 0.0041 percent false-positive rate sounds negligible — until you multiply. At that claimed rate, one million scans of genuinely human writing would produce roughly 41 false accusations; at the older 0.01 percent rate, about 100. Turnitin’s self-reported document-level rate of “under 1 percent” implies up to 10,000 per million. Substack-scale deployment means millions of scans, and each false flag is a human being facing an unfalsifiable charge.
Humanizers punch a hole in the wall
A Notre Dame working paper, “Why AI Detection Fails for Academic Integrity,” found that after AI text was passed through a humanizer tool, the 3.2 model caught it less than 4 percent of the time — while flagging lightly AI-edited human abstracts 64 to 80 percent of the time. That is close to the worst possible combination for academic integrity. An entire industry of rewriting tools — we have reviewed Go Humanize AI and Rehumanize.io — exists to exploit exactly this gap, and The Atlantic reported similar results testing humanized text earlier this year.
Short texts, shifting verdicts
The company itself discloses that accuracy degrades under 100 words, that the median input is only about 350 words, and that an identical passage can receive different verdicts scanned alone versus in context. A pasted excerpt from a novel is precisely the short, decontextualised input where the tool is weakest — and precisely how public accusations happen.
Bias, and who gets accused
“I don’t think AI detectors work,” Sam Illingworth, a professor of critical AI literacy at Edinburgh Napier University, told WIRED. “I think that detectors are prejudiced against certain people.” The company disputes this with internal research, but the optics are stark: all three of publishing’s biggest detection scandals — Shy Girl, Daggermouth and Call Me — involved writers of color. As Regina Brooks, president of the Association of American Literary Agents, notes, whose work gets scanned and publicly accused is a human choice, and human choices carry bias.
The blind spot nobody screenshots
By the company’s own measure, human essays heavily rewritten by AI are still classified as human-written 41.37 percent of the time — the price of erring toward “human” on borderline text. The gold standard, in other words, systematically under-detects the most common real-world use of AI: heavy assistance rather than wholesale generation.
Should You Trust Pangram? A Practical Answer
The evidence supports a narrow conclusion: Pangram is the most accurate AI detector independently tested, and it is still not an oracle. Treat the score as strong probabilistic evidence, never as proof.
| Scenario | How much weight to give the score | Why |
|---|---|---|
| Long document, very high score | High — corroborate, then act | False positives on long text were essentially zero in independent testing |
| Short excerpt, borderline score | Low | Accuracy drops under 100 words; verdicts shift with context |
| A clean “human” verdict | Low to moderate | Humanizer tools evade detection over 96% of the time |
| Grounds for public accusation | Never on the score alone | Even a 1-in-24,000 error rate produces real victims at scale |
If you are a writer
Keep receipts. Drafts, version history, notes and research trails are the evidence that rebuts a false flag — the accusation is statistical, but your paper trail is not. Where platforms disclose scores, assume your work will be scanned and consider disclosing AI assistance up front; Spero’s own framing is that deception, not use, is what ends careers.
If you run a school, publisher or business
Write a policy before you buy a scanner. Decide what score triggers what process, who reviews it, and what corroboration is required — a detection score should open a conversation, not close one, echoing Penguin Random House’s “not determinative” stance. Detection also pairs with provenance: watermarking systems like Google’s SynthID and the EU’s incoming AI content-label rules attack the same problem from the supply side. Our AI governance framework for SMEs covers how to fold both into a workable policy.
FAQ
Is Pangram actually the most accurate AI detector?
On the independent evidence, yes. The NBER benchmark found it was the only detector among four tested to combine near-zero false positives with near-zero false negatives, and the only one meeting a 0.5 percent false-positive policy cap. That is a genuine, measured lead — not marketing.
Can Pangram be fooled?
Yes. Humanizer tools that rewrite AI output evaded the 3.2 model more than 96 percent of the time in Notre Dame’s testing, and writers have deliberately mimicked AI style to trigger false flags. Detection is an arms race, and evasion currently has the cheaper weapons.
What does a Pangram percentage actually mean?
It is a model’s best-guess estimate of how much AI was involved in generating the text — not a measured fact. The tool deliberately errs toward calling borderline text human, which keeps false accusations rare but lets heavily AI-assisted writing pass as human 41.37 percent of the time by the company’s own figures.
Should schools act on a score alone?
No. The company’s disclosed weaknesses — short texts, context-dependent verdicts, humanizer evasion — plus the bias concerns raised by researchers mean a score should trigger review and conversation, never automatic penalties. That is also how Penguin Random House says it treats detection internally.
Who is behind Pangram?
Stanford graduates Max Spero (ex-Google, Nuro) and Bradley Emi (ex-Tesla Autopilot, Absci) founded the company as Checkfor.ai in 2023. It has 24 employees, $13 million in funding led most recently by Menlo Ventures, and counts Substack as its highest-profile integration.
References
Pangram Has Emerged as the Gold Standard of AI Detection. Should You Trust It? — WIRED
Artificial Writing and Automated Detection — NBER Working Paper 34223
Technical Report on the Pangram AI-Generated Text Classifier — arXiv
As AI content floods the internet, Pangram raises $9M to detect it — TechCrunch
Yes, AI detection can be accurate — Pangram
Pangram raises $4M to detect AI-written content — Tech Startups
GPT detectors are biased against non-native English writers — Liang et al.
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.