AI-assisted writing is now the hardest thing on the internet to label, and that is the argument Pangram co-founder and chief executive Max Spero brought to TechCrunch’s Equity podcast on 2 September 2026. The 24-minute episode, produced by Theresa Loconsolo and hosted by Rebecca Bellan, was published under a blunt headline: AI detection is harder than “Real or Fake”. Spero’s point is that the binary verdict everyone wants is the least useful thing a detector can give you.
The framing matters because the stakes have moved out of the classroom. TechCrunch introduced the episode with three examples that have nothing to do with homework: job applications, product reviews and insurance claims. In every one of those, someone is making a decision about a real person based on text they cannot verify. A machine that shouts “fake” at an AI-assisted draft written by a human with a spellchecker does not help that decision. A machine that estimates how much of the text came from a model might.
We covered the company itself last week, when Pangram was being described as the gold standard of AI detection after a run of publishing scandals. This piece is about the argument Spero made on the podcast, the product decisions that follow from it, and what a spectrum-based detector changes for anyone who has to judge submitted text at scale. Our AI models and tools hub tracks these releases as they land. Where sources disagree on a number, we say so rather than picking the one that reads best.
Table of contents
- What Max Spero Actually Said on the Equity Podcast
- Why “Real or Fake” Is the Wrong Question for AI-assisted Work
- How Pangram Measures AI-assisted Text
- What Independent Testing Says About AI-assisted Detection
- Substack Put AI-assisted Scanning in Front of Readers
- The New Image Detector and Why It Is Harder Than Text
- Dead Internet Theory Is Closer Than the Phrase Suggests
- Where AI-assisted Detection Still Fails
- What AI-assisted Detection Means for Your Business
- What We Could Not Verify
- AI-assisted Detection: Common Questions
- References and Further Reading
What Max Spero Actually Said on the Equity Podcast
The episode ran under two headlines on TechCrunch. The video was published as “Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake'”. The audio episode carried a sharper line: “We’re ‘dangerously close’ to dead internet theory, says Pangram’s CEO”.
The line between assisted and generated
Spero’s central claim is that figuring out how much AI went into a piece of writing is both harder and more valuable than labelling it AI or human. A newsletter dictated by its author and tidied by a model is not the same artefact as a newsletter produced whole from a one-line prompt. Both are AI-assisted in some sense. Only one of them is a fake byline.
The dead internet warning
Spero told Bellan that the web is “dangerously close” to dead internet theory becoming reality within a few years — the idea that most of what you read online was written by machines for other machines. He has publicly estimated that AI-generated content already accounts for between 35% and 50% of the internet, and that on X, 29% of long-form posts and 9% of short-form posts are machine-written.
False positives are the real risk
The episode does not treat detection as a solved problem. Spero flagged false positives as the thing that keeps him up, and singled out sensitive images as the case where a wrong accusation does the most damage. He also raised the labour question: entry-level writing jobs shrink, while genuinely human writing becomes more valuable — which only works if you can tell the difference.
| Date | Milestone |
|---|---|
| 2023 | Spero and Bradley Emi found the company as Checkfor.ai |
| 2024 | Renamed Pangram Labs |
| June 2025 | Seed round closes |
| 21 July 2026 | Substack ships reader-facing scanning built on Pangram |
| 29 July 2026 | $9m round led by Menlo Ventures; Pangram 4 and image detection ship |
| 20 August 2026 | Columbia Journalism Review profiles Spero as “the internet’s slop janitor” |
| 2 September 2026 | Spero appears on Equity to argue against the binary verdict |
Why "Real or Fake" Is the Wrong Question for AI-assisted Work
The binary framing survives because it is easy to build a policy around. A journal, a hiring manager or a newsletter platform wants a yes or a no. The problem is that the underlying reality does not come in two states, and forcing it into two states is what produces the false accusations that have made detectors politically radioactive.
A spectrum, not a switch
Consider five ways a paragraph can reach a page. A human types it unaided. A human types it and a model fixes the grammar. A human dictates it and a model shapes it into prose. A human prompts a model and heavily rewrites the output. A model writes it and nobody reads it before publication. Only the last one is a lie about authorship, yet a real-or-fake detector has to sort all five into two bins.
Where the damage lands
The middle three cases are where AI-assisted work lives, and they are also where false positives cluster. Detectors have a documented history of misreading writing by non-native English speakers and neurodivergent writers, both of whom produce text that is more regular and more formulaic than the average native-speaker draft. That is not evidence of a machine. It is evidence of someone being careful.
What a percentage buys you
Pangram’s answer is to return an estimate of how much of a document appears to be model-written rather than a single verdict. A percentage lets an institution set its own threshold, and — more importantly — lets it set different thresholds for different decisions. Rejecting a conference paper and declining a freelance pitch should not use the same bar.
How Pangram Measures AI-assisted Text
Understanding why the company thinks it can do this needs a little detail about the training method, which is unusual among detectors.
Synthetic mirroring
Pangram’s approach is to take a piece of genuinely human writing and have today’s leading models produce a close match on the same subject, in the same format, at the same length. The classifier then learns the difference between the two on near-identical content, rather than learning the much easier and much less useful difference between a Reddit comment and a chatbot answer. It is a natural language processing problem solved by controlling for topic.
Why that matters for false positives
Most detectors fail because they are really measuring predictability. Formulaic human prose looks predictable, so it gets flagged. Training on mirrored pairs strips topic and format out of the signal, which is the mechanism behind Pangram’s headline false-positive claim. Whether the effect survives contact with genuinely unusual writers is a separate question, and one the company’s critics keep asking.
The published false-positive numbers
Pangram 4 is reported at a false-positive rate of 0.0041%, or roughly one document in 24,000; the previous generation sat at about one in 10,000. The image model, still in research preview, is much looser. On the ReLAION set of 10,000 pre-2022 images it returns 0.16%, about one in 625, and on its own internal benchmark it reports 0.40% on clean data. On a 2,000-image subset of WikiArt it flagged nothing at all.
What Independent Testing Says About AI-assisted Detection
Vendor benchmarks are vendor benchmarks. The reason Pangram is taken seriously is that outside groups keep arriving at similar conclusions, which is rare in this category.
The peer-reviewed evaluations
Nature reported in 2026 that GPTZero claims 99% accuracy and Pangram claims 99.98%, and noted that the American Association for Cancer Research now runs Pangram over peer-review reports because of undisclosed model use. A June 2026 peer-reviewed study from Vrije Universiteit Brussel tested four detectors and found that only Pangram produced satisfactory results. A 2025 University of Chicago Becker Friedman Institute working paper found essentially zero false positives and false negatives on medium-to-long passages.
Deployment at scale
Two deployment numbers give a sense of volume. Pangram says one in eight biomedical articles last year contained some AI-generated text. In June 2026, NeurIPS announced that it rejected 18% of submissions after screening them with Pangram. A University of Maryland analysis found AI text in more than 9% of 186,000 articles drawn from 1,500 newspapers in summer 2025.
The dissent
The counter-case is institutional rather than statistical. MIT and the University of Waterloo both advise against treating any detector output as a definitive verdict, and several universities have withdrawn detection tools entirely over unreliability and bias against non-native English speakers. Technology journalist Taylor Lorenz called Pangram the best tool on the market in April, which is praise with a ceiling: best in a category people distrust.
| Evaluation | Result reported |
|---|---|
| Vrije Universiteit Brussel, June 2026 | Only detector of four tested judged satisfactory |
| Becker Friedman Institute, 2025 | Essentially zero false positives and negatives on long passages |
| NTIRE 2026 (images) | 99.999% AUROC |
| Mirage-Test (images) | 99.66% accuracy |
| Synthbuster (images) | 98.49% macro accuracy |
| Community Forensics (images) | 97.29% accuracy |
Substack Put AI-assisted Scanning in Front of Readers
The clearest test of Spero’s thesis is not a benchmark. It is what happened when the spectrum idea met a live writing platform with a large, opinionated user base.
What actually shipped
Substack rolled the feature out on 21 July 2026, on the web first, with mobile promised later. Readers and writers can scan any post longer than 100 words and get an estimate of model involvement. Nothing is blocked, flagged or punished by default — the scan only runs when a person asks for it, writers can disable it on individual posts, and anyone can report a result they believe is wrong.
The disclosure field
Alongside the scanner, Substack added a voluntary “how I make this” field, where a writer can describe their process: dictation, an editing pass, research help, or nothing at all. That field is the more interesting half of the release, because it is the platform conceding that AI-assisted work is normal and that the useful signal is disclosure rather than detection.
The backlash
Writers did not take it well. The response included accusations of a witch hunt and renewed complaints about false positives hitting non-native speakers. Chief executive Chris Best defended it on trust grounds: “When readers have to wonder if what they’re reading is real, it undermines trust in authorship and threatens the livelihood of writers.” Spero’s version is shorter — “If you’re not going to bother writing your newsletter, why should I read it?”
The New Image Detector and Why It Is Harder Than Text
Pangram shipped image detection as a research preview on 29 July 2026, the same day as the funding round. The gap between its text and image accuracy is the clearest evidence for Spero’s own argument that this problem is not one problem.
What it covers
The model targets outputs from OpenAI’s GPT Image, Google’s Gemini Nano Banana, Midjourney, FLUX, Grok Imagine, Riverflow v2, Flux 2 Pro and Bytedance Seedream, plus video from Kling, Seedance, Veo, Wan and Grok Image Video. On a 1,130-image internal benchmark it reports 100% accuracy on clean images and 99.03% after augmentation, against 99.13% for the next closest competitor and 98% for the previous best published result.
What it cannot do
The limits are published and they are significant. It will not process images under 512 by 512 pixels. It does not detect deepfakes or face swaps — a notable exclusion given that face swaps are the format doing the most real-world harm. It refuses NSFW content, and API access is invitation-only during the preview.
Video is where the error rate lives
Across nine video models the tool reports 98.91% accuracy on clean frames and 97.08% after augmentation, but the per-model range runs from 85% to 100%. Read as error rates, that is 1.09% on clean video and 2.92% after augmentation, with a worst case of 15% on whichever model it handles least well.
Dead Internet Theory Is Closer Than the Phrase Suggests
Spero’s “dangerously close” remark sounds like founder marketing until you line it up against measurements taken by people with no detector to sell.
What the crawl data shows
A Stanford analysis of Internet Archive captures from late 2022 to mid-2025 found that by May 2025, 35.3% of newly published websites were built with AI assistance, and 17.6% were entirely machine-generated. That is the supply side of the problem: not slop hidden in the corners, but a third of new pages.
What the traffic data shows
The demand side is worse. Cloudflare put bots at 57.5% of webpage requests as of June 2026. Imperva’s 2026 report has automated traffic above 53% against 47% human, up from 51% in 2024. HUMAN Security reported that traffic from agents that actually click links and fill in forms grew 7,851% year on year. We looked at what that does to publishers in our piece on AI crawlers eating website traffic.
Why detection follows from that
If most requests are machine-made and a third of new pages are AI-assisted, then provenance stops being an academic question and becomes basic hygiene. Platforms have started acting on it: Instagram is cracking down on AI accounts posing as humans, and the EU has made AI content labels compulsory for authentic-looking material.
Where AI-assisted Detection Still Fails
Spero’s own framing invites a harder question than the marketing does: if the answer is a percentage, what is that percentage actually measuring, and when is it wrong?
Humanisers and light edits
The uncomfortable finding across the research is that the two ends behave differently. Heavily rewritten model output can slip past detectors, while lightly edited human text can trip them. That asymmetry is exactly backwards from what a fair system would do, and it is the strongest technical argument for reading a score as evidence rather than a verdict.
Short text
Every strong result in this category comes with a length qualifier. The Becker Friedman finding is explicitly about medium-to-long passages, and Substack’s scanner will not touch anything under 100 words. Comments, headlines, product reviews and chat messages — a large share of the AI-assisted content people actually encounter — sit below that floor.
The disclosure gap
A detector cannot tell you intent. It cannot distinguish a writer who used a model and said so from one who used a model and hid it, which is the distinction that actually matters to a reader. That is why Substack’s “how I make this” field and watermarking schemes like Google’s SynthID verification sit alongside detection rather than being replaced by it.
What AI-assisted Detection Means for Your Business
If you commission writing, screen applicants, moderate a platform or review submissions, Spero’s argument has a practical shape. Treat the output as a probability, not a proof.
Set thresholds by consequence
The cost of a false positive is not constant. Declining to publish a guest post is recoverable; firing a contractor or rejecting a candidate is not. Set a high threshold where the consequence is severe, a lower one where it merely triggers a conversation, and never automate the severe end.
Ask before you accuse
The cheapest control is a disclosure policy. Tell writers and applicants what level of AI-assisted work is acceptable before they submit, and most of the problem disappears without a detector being involved. We went through the signals people can spot themselves in our guide to detecting AI writing and its limits.
Keep a human in the loop
Every credible institution in this space, Pangram included, says the same thing: a score is an input to a judgement, not the judgement. Log the score, log the decision, and let the person accused respond. The same principle governs AI-generated code review, where the Debian project chose provenance rules over an outright ban.
| Decision | Suggested handling |
|---|---|
| Publishing a guest article | Score as a prompt to ask about process, not to reject |
| Screening a job application | Never automated; disclosure policy stated up front |
| Moderating user reviews | Below the length floor; use behavioural signals instead |
| Reviewing an academic submission | High threshold, long passages, right of reply |
| Paying a freelance writer | Contract terms on AI-assisted work beat any detector |
What We Could Not Verify
Some of this reporting does not reconcile cleanly, and it is worth being explicit about which parts.
The funding history
Sources disagree on the size of the 2025 seed round. One account puts it at $2.7m and another at $4m, both against the same $9m Menlo Ventures round in July 2026 and the same roughly $13m total. We have not seen a filing that settles it.
The accuracy headline
The 99.98% figure Nature attributes to Pangram is a vendor claim, and the 0.0041% false-positive rate comes from internal benchmarks. The independent studies are supportive but were run on the datasets their authors chose. None of them tell you how the model behaves on your submissions.
The podcast itself
TechCrunch published the episode description and the video, but no transcript. Everything attributed to Spero above comes from the published episode summaries and his prior on-record statements, not from a verbatim transcript we have checked line by line.
AI-assisted Detection: Common Questions
What did Max Spero say on the Equity podcast?
That AI detection is harder than a real-or-fake verdict, and that estimating how much AI went into a document is both more difficult and more useful than a binary label. He also warned that the web is “dangerously close” to dead internet theory.
Is AI-assisted writing the same as AI-generated writing?
No, and that distinction is the point of the episode. AI-assisted covers everything from a grammar fix to a heavy rewrite of model output. AI-generated means nobody meaningfully wrote it. Only the second is a false byline.
How accurate is Pangram?
Pangram claims 99.98% accuracy and a false-positive rate of 0.0041% on text, roughly one in 24,000 documents. Independent studies from Vrije Universiteit Brussel and the Becker Friedman Institute broadly support the text results. Image detection is weaker, at 0.16% false positives on the ReLAION set.
Can Pangram detect AI images and video?
Yes, in research preview since 29 July 2026, covering GPT Image, Nano Banana, Midjourney, FLUX and others plus several video models. It cannot handle images below 512 by 512 pixels, deepfakes or face swaps.
What is Substack doing with Pangram?
Since 21 July 2026, readers can run an opt-in scan on posts over 100 words. Nothing is blocked automatically, writers can disable it per post, and a voluntary “how I make this” field lets authors describe their process.
Why do AI detectors flag non-native English speakers?
Because most detectors are measuring predictability, and careful, formulaic prose is predictable. Pangram trains on mirrored human and model versions of the same text to control for that, but no detector has eliminated the bias.
Should I use a detector on job applications?
Not as an automated filter. Use a stated disclosure policy instead, treat any score as one input among several, and give the applicant a chance to respond before a decision is made.
References and Further Reading
TechCrunch on Why AI Detection Is Harder Than Real or Fake
TechCrunch Equity on Dead Internet Theory
Columbia Journalism Review on Pangram’s Slop-Free Future
Nature on How Good AI Detection Tools Have Become
Pangram on Its Image Detection Research Preview
Pangram on Third-Party Evaluations
Pangram on False Positives in AI Detectors
Substack on How Writers Reacted to Its Transparency Tools
The Pangram 4 Technical Report
International Journal for Educational Integrity on Detector Reliability
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.