AI detectors are having their most public stress test yet. On the evening of Monday 21 September 2026, an anonymous account on X called Balance ton Claude posted screenshots claiming that Pangram, the best-regarded of the tools, had scored passages of C’était ça ou mourir as “100% AI”. The novel is the best-selling debut of Canadian-Haitian writer Thélyson Orélien, and it is in the running for the Prix Goncourt. Within a day the post had been viewed more than three million times, and a prize contender was on trial by software.

Orélien, 38, denies using AI. His French publisher, Grasset, calls the affair a hate campaign. And when AFP, franceinfo, Le Devoir and Radio-Canada ran the book through AI detectors of their own, the answers ranged from “very likely” human to 100% machine-written. That spread is the real story. It is the clearest natural experiment so far on a question that every teacher, editor, publisher and hiring manager now faces: how reliable are AI detectors, and what can a score actually prove?

This article works through the evidence. It sets out what each tester found in the Orélien case, how AI detectors work, what the independent benchmarks say, where the tools break down, and the arithmetic that decides whether a flag means anything. It builds on our earlier pieces on whether you can teach yourself to detect AI writing and on Pangram’s rise to gold-standard status. It ends with a practical policy for organisations that want to use AI detectors without wronging anyone.

One Novel, Many Verdicts: The Orélien Case and AI Detectors

ai detectors spotting ai writing how reliable b traffic light with three lamps

How the accusation started

C’était ça ou mourir (“It was that or die”) follows a Haitian migrant travelling to Canada. It is published by Grasset in France and by Boréal in Quebec. It won the Prix Fnac 2026, has been sold to publishers in more than 20 countries, and Le Devoir lists it among the finalists for the Goncourt, Médicis, Femina and Renaudot prizes this autumn. Doubts predated the X post. Podcaster Léa Bory suggested on Torchon that it read like AI because of repetitive sentences and a heavy use of analogies, and critic Fabrice Colin wrote in Le Canard enchaîné that he had “vainly looked for a unique voice”.

Balance ton Claude, whose name nods to Anthropic’s Claude model, had fewer than 100 followers before the post, according to franceinfo, and several thousand 24 hours later. “We tested several passages, at various points in the manuscript, for almost always the same result: 100% AI,” it wrote (our translation). It later said it had run four texts Orélien wrote in 2014 through Pangram and got 0% AI for each. It then urged people to go easy on the author, since there was a “non-zero chance” he was acting in good faith.

What the testers found with AI detectors

Four newsrooms and one poet then ran their own tests. The table below gathers every published result we could verify. None of the testers used identical text, tool settings or sample length, which is part of why the verdicts diverge so sharply.

TesterToolWhat was testedResult
Balance ton Claude (X)PangramSeveral excerpts, including the first two chaptersAlmost always 100% AI
AFPSeveral AI detectorsExcerpts of the bookFrom “very likely” human to 100% AI
franceinfoPangram, free versionExcerpts of varying length, dialogue and description, some mixing French with English and Spanish100% AI in every case
franceinfoGPTZeroOne excerpt87% chance of AI
Le DevoirInVID-WeVerify, co-created by AFPThe passages cited on X“Very likely” written by a human
Le DevoirSeven other AI detectorsThe same passagesOne sided with Pangram, two said human, four leaned human without committing
Radio-CanadaPangramThe full ebook94% of the text flagged as AI-generated
Poet Emmanuel DerapsPangramOne excerpt, then the same excerpt with about 20 words changed100% AI, then 100% human

Read down the right-hand column and the problem is plain. Eight tests, at least a dozen AI detectors, and verdicts that span the whole scale. Some of that spread is noise from weak tools. Some of it is how the strongest tool is designed to behave. Separating the two is the job of the rest of this article.

The controls that complicate the story

The more careful testers also ran controls: texts whose authorship is not in doubt. Those results point the other way from the author’s defence, at least for Pangram on these samples.

TesterControl text run through PangramResult
Balance ton ClaudeFour Orélien texts from 20140% AI each
franceinfoOrélien’s 2015 book Le temps qui reste and a 2014 HuffPost article100% human
franceinfoExtracts of four other autumn 2026 novels, by Ananda Devi, Philippe Jaenada, Amélie Nothomb and Lucie Rico99% to 100% human
Radio-CanadaSix other Goncourt-listed novels available digitally in Quebec100% human
Radio-CanadaEight books by Haitian-born authors, including Lyonel Trouillot, Marie-Célie Agnant and Jacques Roumain, plus one by Blaise Ndala100% human
Radio-CanadaWhole chapters of the Bible and nine French classics, from Maupassant to Foucault100% human

These controls answer the two defences raised most often. Orélien told Libération the software was “used to the French literary tradition and classical authors”. Grasset told BFMTV that detection tools “fabulate that the Bible was written artificially”. Marzena Karpinska of Simon Fraser University told Radio-Canada that some AI detectors do flag very short excerpts of much-quoted old texts, because those texts saturate the training data of language models. She said Pangram is not one of them, and that the novel still scored as AI when translated into English, which human texts translated the same way do not.

Radio-Canada later added a disclosure. Karpinska co-wrote a research paper with several Pangram employees, on AI-generated articles in US media, and Pangram donated credits for that work. She says she has no financial or professional tie to the company.

Where the case stands

None of this proves authorship, and Karpinska was careful to say so. “We cannot say with 100% certainty that it is AI,” she told Radio-Canada. “We can say there is a very high probability that this text was produced with significant AI assistance.” Entrepreneur Benoît Raphaël, who ran his own Pangram tests, told Nouvel Obs that a rating of “around 100% AI, or even 90%, is a pretty telling sign.”

On the other side, Orélien told La Presse that his first manuscript was finished in 2019, before ChatGPT existed. Boréal says the text matured through many rounds with his editor, “as the successive versions of the manuscript attest”. That version history is exactly the kind of evidence that could settle the question, and at the time of writing it had not been published. The sociologist Michel Lacroix told Le Devoir that a single nervous jury member could now sink the book’s prize chances, whatever the truth.

How AI Detectors Actually Work

ai detectors spotting ai writing how reliable c metronome with its arm swung aside

Classifiers trained on paired text

Most commercial AI detectors are classifiers. They are trained on large collections of human writing alongside machine-generated text, and they learn the statistical features that separate the two. AFP’s summary is fair: the tools hold “large troves of AI-generated text alongside content written by humans” and compare the two.

Pangram’s twist, as Karpinska described it, is to pair the data. Its makers took millions of human-written texts, asked several AI models to rewrite the same material, and trained on those matched pairs. The model therefore learns how a human and a machine version of the same content differ, rather than learning topic or genre by accident. That design is one reason Pangram has pulled ahead of older AI detectors in independent tests.

Statistical scores and “linguistic tics”

Older and open-source tools lean on statistical scores such as perplexity: how predictable each next word looks to a language model. Machine text tends to be more predictable than human text, but so does plain, careful, conventional human prose. That overlap is the root of most false positives.

Thierry Poibeau, a researcher at France’s CNRS and a specialist in natural language processing, told AFP that the programmes look for “linguistic tics”. “The classic example is triadic phrasing (using three clauses), or em dashes,” he said. Those tells were common in early chatbot output because the models had been trained heavily on American scientific writing and American novels. Recurring words and phrases, and monotonous sentence structures, are further warning signs.

What the percentage on AI detectors means

A score such as “94% AI” is easy to misread. Radio-Canada’s 94% is the share of the book’s text that Pangram attributed to AI. It is not a 94% probability that the author cheated. GPTZero’s “87% chance” in franceinfo’s test is a different kind of number: a probability for one excerpt. Some AI detectors report a “human” percentage instead of an “AI” one, which is why the Authors Guild had to normalise every result before comparing tools.

The industry is slowly moving from binary verdicts towards estimates of how much AI went into a document, as we covered in our interview piece on Max Spero’s case for AI-assisted scoring. That is more honest, and harder to read at a glance. A reader who sees “94%” and thinks “94% guilty” has already made the most common mistake with AI detectors.

Why AI Detectors Disagree on the Same Text

ai detectors spotting ai writing how reliable d two dice showing pips

Different tools, very different quality

The first reason is simply that AI detectors are not equally good. In the Authors Guild’s test, covered below, one consumer tool flagged every pre-AI human article as mostly machine-written. When a newsroom runs a passage through a dozen tools of mixed quality, the spread it reports is partly a spread in quality. Averaging a good detector with a bad one tells you very little.

Excerpts, length and free tiers

Short samples give a classifier less to go on. Turnitin raised its minimum from 150 to 300 words for exactly this reason, saying accuracy improves “with a little more text”. The University of Chicago benchmark found Pangram robust even on “stubs” of 50 words or fewer, but most AI detectors are not. The Orélien testers worked with excerpts of different lengths and, in franceinfo’s case, the free version of Pangram. Radio-Canada’s full-text run is the more informative of the tests.

Small edits, big swings

The poet Emmanuel Deraps told Le Devoir he replaced about ten words in a passage, deleted four or five and added four or five more. The first version scored 100% AI; the edited one scored 100% human. To many readers that looked like proof that AI detectors are random.

Karpinska’s reading is different. Light editing is how people “humanise” AI text to escape detection, she said, and it is easier to fool Pangram this way “because the tool is optimised to offer a minimal rate of false positives, while increasing the chance of false negatives”. In other words, a flip from AI to human after a light edit is the tool behaving as designed: it would rather miss AI text than accuse a human. The flip the other way, with human text scored as AI, is the rarer and more serious error.

Language, genre and literary style

Poibeau’s warning to AFP applies directly to fiction: “The tool may pick up something particular in the writing, something repetitive, but that might be exactly what constitutes a writer’s style.” Boréal made the same point in its statement, saying the tool’s reliability “has never, to our knowledge, been established on French-language novelistic prose”.

Karpinska told Radio-Canada that Pangram’s false-positive rate on French text is 0.0026%, and essentially zero on English literary text, but conceded there is no data specific to French literary prose. Her view is that fiction is easier, not harder: literary texts vary far more in form than news articles do, while AI tends to reproduce the same patterns. Filip Van Droogenbroeck of the Vrije Universiteit Brussel told Libération that Pangram deliberately trained on multilingual data, including under-represented languages.

Writers are not reassured. The French novelist Baptiste Beaulieu told Le Devoir he has always loved ternary rhythms and uses them everywhere. “So what do I do now?” he asked. “Change the way I write so that a piece of software agrees to believe I am human?”

What Independent Tests Say About AI Detectors

ai detectors spotting ai writing how reliable e weathervane with a flat arrow

Vendor marketing is not evidence, so the useful question is what researchers with no stake in the outcome have found. The table summarises the six studies most often cited in this debate, from newest to oldest.

StudyTools testedSampleHeadline finding
Vrije Universiteit Brussel, IJEI, June 2026GPTZero, Pangram, Copyleaks, Turnitin160 synthetic academic papers; 1,163 master’s thesesPangram best; the others underestimated AI content; all four identified fully human texts
Authors Guild, May 2026Pangram, Originality.ai, Grammarly, ZeroGPT, SidekickerTen Guild articles from 2022 or earlierPangram 0% AI on all ten; Sidekicker 71% to 100% AI on all ten
University of Chicago, NBER w34223, 2025Pangram, Originality.ai, GPTZero, RoBERTa1,992 human and 1,992 AI texts across genres, lengths and modelsPangram near-zero error rates; the only tool to meet a 0.5% false-positive cap
Russell, Karpinska and Iyyer, ACL 2025Pangram, GPTZero, Fast-DetectGPT, Binoculars, RADAR, human experts300 articlesPangram caught 98.0% of AI articles; RADAR caught 15.3%
Liang et al., Stanford, Patterns, July 2023Seven early AI detectors91 TOEFL essays plus US eighth-grade essays61% of non-native essays falsely flagged, on average
OpenAI, January to July 2023OpenAI’s own classifierOpenAI’s challenge setCaught 26% of AI text and falsely flagged 9% of human text; withdrawn

The University of Chicago benchmark

The paper AFP cites, “Artificial Writing and Automated Detection” by Brian Jabarian and Alex Imas, is the most rigorous head-to-head of commercial AI detectors so far. It tested three commercial tools and one open-source model across a large corpus spanning genres, lengths and generating models. The commercial tools beat the open-source one, and Pangram achieved “near-zero” false-negative and false-positive rates that held across models, thresholds, ultra-short passages and “humanizer” tools.

Its most useful contribution is a framework rather than a ranking. The authors let a decision-maker set a “policy cap”, a maximum tolerable error rate that does not depend on the tool. Under a strict cap of 0.5% false positives, Pangram was the only one of the four AI detectors that met it without sacrificing accuracy.

The Vrije Universiteit Brussel study

Published in the International Journal for Educational Integrity in June 2026, the VUB study built 160 academic papers with known ground truth: fully human, fully AI, hybrid, and AI text “humanised” by a prompt designed to resemble what a student might do. Pangram consistently outperformed GPTZero, Copyleaks and Turnitin. The other three “significantly underestimated” AI content, especially from the newest model.

The authors then ran Pangram over 1,163 real master’s theses from the 2024-2025 academic year. It flagged 45.5% of them, typically at low to moderate levels of AI-associated text. False positives were rare across all tools in the controlled set, yet the authors still concluded that AI detectors “should not be used as sole evidence in high-stakes decision-making”.

The Authors Guild test of AI detectors

In May 2026 the Authors Guild, the largest professional body for US writers, ran ten of its own articles, all published in 2022 or earlier, through five tools. Any accurate tool should have scored every one as human. The chart shows the highest AI score each tool gave to any of the ten.

Highest AI score given to a pre-2023 human article (Authors Guild, May 2026)
Sidekicker 100%
ZeroGPT 76%
Grammarly 9%
Originality.ai 1%
Pangram 0%

The detail is worse than the maximums. ZeroGPT scored the Guild’s obituary of Joan Didion at 66% AI and its congratulations to Louise Erdrich on her Pulitzer at 76%, with no pattern the Guild could explain. Sidekicker scored every article between 71% and 100%. The Guild’s verdict: “some commonly used consumer-facing AI detection tools are wildly inaccurate”, and even the best “should never be the sole basis for any decision”.

The expert-reader study

Jenna Russell, Marzena Karpinska and Mohit Iyyer tested five AI detectors against 300 articles at ACL 2025, and compared them with human readers. At a fixed 2% false-positive rate, Pangram caught 98.0% of the AI-written articles, GPTZero 85.3%, Fast-DetectGPT 80.0%, Binoculars 66.7% and RADAR 15.3%. A panel of frequent chatbot users, voting by majority, missed just one AI article in 300. This is the “2024 test” Karpinska described to Radio-Canada, and it is where Pangram’s reputation began.

The older warnings

Two 2023 results still shape the public view. OpenAI launched its own classifier on 31 January 2023, reported that it caught only 26% of AI-written text while falsely flagging 9% of human text, and withdrew it on 20 July 2023 “due to its low rate of accuracy”. That same year, Weixin Liang and colleagues at Stanford found seven popular AI detectors wrongly flagged 61% of TOEFL essays by non-native English speakers on average, while passing more than 90% of essays by US eighth-graders.

Put together, the pattern is clear. Since 2023 the field has split. A small number of AI detectors, with Pangram the most consistent, now post very low false-positive rates on long English text in independent tests. Many widely used consumer tools do not. So “are AI detectors reliable?” has no single answer. It depends on which tool, which text and which error you care about.

The Arithmetic of False Positives

ai detectors spotting ai writing how reliable f gavel resting on a sound block

Two errors with two different harms

The Chicago paper frames the problem in two numbers. The false-negative rate is the share of AI-generated text a tool wrongly passes as human. The false-positive rate is the share of human-written text it wrongly flags as AI. A school worried about cheating cares about the first. A novelist accused of cheating cares about the second. Most AI detectors let the operator trade one against the other by moving a threshold, and Pangram, by Karpinska’s account, is deliberately tuned to keep false positives minimal at the cost of missing more AI text.

Vendor claims, and why they are hard to check

AFP made a blunt observation: many tools publish percentage figures about their accuracy, but the figures are impossible to verify and do not account for the risk of a false positive. The published numbers that do exist come from very different tests.

Turnitin, after running 800,000 pre-ChatGPT academic papers through its system, said its document-level false-positive rate is under 1% for documents where it detects more than 20% AI writing, and that its sentence-level rate is about 4%. Scores below 20% now carry an asterisk because they are less reliable. Pangram’s own site put its overall rate at about 1 in 10,000 (0.01%) as of 15 September, and 0.003% for books, according to Radio-Canada. Earlier this year the company reported 0.0041% for its Pangram 4 model. Vendor figures move, and they are measured on the vendor’s own test sets.

To make those rates concrete, multiply them out. The chart shows how many human-written documents out of 10,000 each stated false-positive rate would wrongly flag.

Human documents wrongly flagged per 10,000, at each stated false-positive rate
OpenAI classifier, 9% (withdrawn 2023) 900
Turnitin document level, under 1% up to 100
Pangram overall, about 0.01% about 1
Pangram on books, 0.003% 0.3

The gap between the best and worst AI detectors is three orders of magnitude. At Turnitin’s upper bound, a university that screens 10,000 honest essays could still produce up to 100 false flags. Behind each, as Turnitin itself wrote, “is a real student who may have put real effort into their original work”.

Base rates decide what a flag means

A false-positive rate alone does not tell you how likely a flagged document is to be innocent. That depends on how common AI writing is in the pile you are screening. The worked example below takes a hypothetical detector that catches 90% of AI text and falsely flags 1% of human text, Turnitin’s stated upper bound, and applies it to 10,000 documents at three levels of AI use. The last row repeats the lowest level with a 0.01% false-positive rate.

Share of AI-written documentsAI documents caughtHuman documents flaggedFlags that are wrong
2% (200 of 10,000)1809835.3%
10% (1,000 of 10,000)900909.1%
30% (3,000 of 10,000)2,700702.5%
2%, with a 0.01% false-positive rate180about 10.5%

The lesson is uncomfortable. Where AI use is rare, as it should be among shortlisted literary novels, even a 1% false-positive rate means more than a third of flags point at innocent writers. Only a far lower rate makes a flag meaningful in a low-prevalence pile. That is why the choice of tool matters so much, and why the same score means different things in a first-year essay queue and a prize shortlist.

Share of flags that are wrong in the worked example
2% AI use, 1% false-positive rate 35.3%
10% AI use, 1% false-positive rate 9.1%
30% AI use, 1% false-positive rate 2.5%
2% AI use, 0.01% false-positive rate 0.5%

Where AI Detectors Are Weakest

Short texts, openings and endings

Every serious study agrees that length helps. Turnitin found more false positives in the first and last few sentences of a document, typically the introduction and conclusion, and changed how it aggregates them. OpenAI said its classifier was “very unreliable” below 1,000 characters. Excerpts pasted into a free web form, the way most people test a suspicion, are the least reliable input AI detectors see.

Humanisers and light edits

An entire industry sells rewriting tools that exist to beat AI detectors. We have reviewed two of them, Go Humanize AI and Rehumanize.io. The Chicago benchmark found Pangram robust to the humanisers it tested, while the VUB study found the other three tools badly underestimated humanised text. The Deraps experiment shows the limit: a person making careful edits by hand can push a passage below the threshold of a tool tuned to avoid false positives.

Non-native and highly polished writers

The Stanford result still defines the bias debate. Early AI detectors punished the simpler, more predictable phrasing typical of people writing in a second language. The Authors Guild describes a mirror-image problem at the other end of the skill scale: every large language model was trained on polished, edited prose, so “the more refined and controlled a writer’s style, the more it may resemble the output these tools are designed to flag”. Neurodivergent students and writers with a strongly patterned voice raise the same concern.

The moving target

Poibeau calls it the “moving target problem”: AI programmes are “constantly evolving”, so a tool trained on last year’s models may miss this year’s. The Authors Guild makes the same point about the detectors themselves, whose accuracy “at any given moment cannot be assumed”. A result from 2023, good or bad, says little about a tool in 2026, and a result from today may not hold after the next model release.

Weak spotMain error it causesEvidence
Short excerptsBoth errors riseTurnitin raised its minimum to 300 words; OpenAI warned below 1,000 characters
Light human editsFalse negativesAbout 20 changed words flipped a Pangram result from 100% AI to 100% human
Humaniser toolsFalse negativesVUB: three of four tools underestimated humanised text
Non-native writersFalse positivesStanford: 61% of TOEFL essays flagged by early tools
New modelsFalse negativesVUB: tools other than Pangram struggled most with the newest model

The Human Cost of a Wrong Call

For authors and publishers

Orélien is not the first writer caught in this. In March, Hachette withdrew the horror novel Shy Girl days before its planned US release after Pangram identified the text as AI-written, Radio-Canada reports. In May, one of the five winning stories of the 2026 Commonwealth Short Story Prize, “The Serpent in the Grove” by Jamir Nazir, scored 100% AI on Pangram. The Commonwealth Foundation and Granta stood by it, noting that AI detectors are imperfect, while the Foundation reviews its selection process.

The Authors Guild’s advice to publishers is firm. They should disclose their methods, keep up with the tools’ accuracy, and give authors “a fair and full opportunity to defend themselves”. They “should never cancel a contract or pull a book on the basis of accusations or use of these tools alone”, which the Guild says would be a material breach of contract.

For students and teachers

Education is where AI detectors are used most. By May 2023, Turnitin had already processed 38.5 million submissions for AI writing. The VUB theses result shows what that looks like in practice: nearly half of real master’s theses carried some flag, mostly at low levels. Each of those flags is a conversation a supervisor has to handle well, and a student who may be innocent.

For businesses, hiring and reviews

The Chicago authors open their paper with business uses: checking that product reviews were written by real customers, not only that coursework was done by students. The same logic now reaches cover letters, supplier proposals, commissioned copy and customer complaints. Pangram’s chief executive, Max Spero, told The Atlantic that the tool should not be the ultimate arbiter, but a starting point for a more in-depth investigation. Organisations using AI detectors for screening should take the vendor at its word.

How to Use AI Detectors Responsibly

The evidence supports a narrow, useful role for AI detectors: triage. The table sets out what a detector score can reasonably support in the situations where organisations meet it most.

SituationReasonable use of a scoreUnreasonable use
Large inbound queue (reviews, applications)Prioritise items for human reviewReject flagged items automatically
Academic submissionOpen a conversation and ask for draftsSole basis for a misconduct finding
Manuscript or commissioned copyRequest version history and process evidenceCancel a contract or pull a book
Text under 300 wordsTreat as a weak signal at mostAny decision at all
Second-language or highly polished writerWeigh against known bias and contextTreat a high score as proof

Treat a score as a lead, not a verdict

Every independent study reviewed here, and the leading vendor itself, says the same thing: a score opens an inquiry and never closes one. Write that into the process, so that no adverse decision can be taken on a detector result without a named person reviewing the wider evidence.

Choose AI detectors on independent evidence

Pick a tool on third-party benchmarks, not the accuracy figure on its landing page. On today’s evidence, that means a tool with a published, independently tested false-positive rate on text like yours. Then test it on your own material. Run a sample of documents you know are human, such as work written before 2022, and a sample you know is AI-generated. If it fails your own controls, its marketing figures are irrelevant.

Test the whole document and run controls

Radio-Canada’s approach is the model. It tested the full text rather than excerpts, then ran the same tool over comparable books whose authorship is not in doubt. When a result matters, do both. A high score on a full document, set against clean scores on the same writer’s earlier work and on comparable texts, is far stronger evidence than one pasted paragraph.

Ask for process evidence

The best evidence of authorship is usually not in the text at all. Version history in a word processor, dated drafts, notes, research trails and an editor’s correspondence all show how a piece came to exist. Ask for them early, and ask neutrally. Orélien’s publishers say such a record exists for his novel, and producing it would say more than any of the AI detectors have.

Write a detection policy

Decide in advance which uses of AI are allowed, which tool you use, what threshold triggers review, who reviews, and how people can respond. Our AI governance framework for SMEs covers the wider policy structure, and our AI strategy team can help you set one up. A written policy protects the organisation as well as the people it screens.

Beyond AI Detectors: Provenance, Watermarks and Disclosure

Watermarks at the point of generation

AI detectors guess after the fact. Watermarking marks text as it is generated, by nudging a model’s word choices into a pattern that can be tested for later. Google has open-sourced its SynthID text method, as we explained in our guide to verifying AI content with SynthID. Watermarks only work for models that apply them, and paraphrasing weakens them, but a positive result is far stronger evidence than a statistical guess.

Labelling duties

Regulators are moving the burden onto providers. Since 2 August 2026, Article 50 of the EU AI Act has required providers to mark synthetic output in a machine-readable way, and deployers to disclose deepfakes and some AI-generated text. We covered the details, and the risks of labelling, in our piece on the EU’s AI content label rules. None of this helps with a novel written by a person using a chatbot privately, which is exactly the Orélien scenario.

Disclosure as a norm

Even Balance ton Claude ended on disclosure. It said it flagged the novel because it was successful, that it was “far from the only one”, and called on authors who use AI to say so rather than betray their readers. The Authors Guild already runs a Human Authored certification for books, and publisher contracts increasingly ask about AI use. They do not replace AI detectors, but they change the question from “can we catch it?” to “did you tell us?”

Verdict: How Reliable Are AI Detectors in 2026?

What the evidence supports

The best AI detectors are now far better than their reputation. On long English text, independent studies in the US and Belgium found Pangram’s false-positive rate close to zero, and several found it reliably catches unedited AI output. In the Orélien case its result was consistent across excerpts, the full book and translation, while its controls on comparable human books came back clean.

What it does not support

The typical free tool is still unreliable, and even the best can be pushed under its threshold by careful editing. No study has established error rates for French literary prose. Scores are routinely misread, excerpts are routinely used in place of full texts, and the arithmetic of base rates means a rare false positive still lands on real people. A score, however high, is not proof of how a text was written.

The bottom line

So, how reliable are AI detectors? The good ones are reliable enough to justify a closer look, and not reliable enough to justify a verdict on their own. That is the same answer researchers, the Authors Guild and Pangram’s own chief executive give. The Orélien affair will be settled, if at all, by drafts and version history, not by another screenshot.

Frequently Asked Questions

Are AI detectors accurate?

It depends heavily on the tool. Independent tests from the University of Chicago and the Vrije Universiteit Brussel found Pangram’s error rates near zero on long English text. The Authors Guild found some consumer tools flagging every human article it tested as mostly AI, with scores up to 100%.

Why did AI detectors give different results for the same novel?

The testers used different tools, different excerpts and different versions. Tool quality varies widely, short excerpts are less reliable, and scores measure different things: a share of text, a probability or a “human” percentage. A light edit of about 20 words also flipped one Pangram result.

Can AI detectors be fooled?

Yes. Humaniser tools and careful hand edits can push AI text below a detector’s threshold, especially for tools tuned to minimise false positives. Evading detection is easier than being falsely accused by the best tools, though weaker tools do falsely accuse human writers.

Do AI detectors discriminate against non-native English speakers?

Early ones did. A 2023 Stanford study found seven popular tools falsely flagged 61% of TOEFL essays on average. Newer tools report much lower rates, but the Authors Guild warns that highly polished human prose can also resemble AI output.

Should a publisher, school or employer act on a detector score alone?

No. Independent researchers, the Authors Guild and Pangram’s chief executive all say a score should start an investigation, not end one. Ask for drafts and version history, test full documents, and run controls before reaching any conclusion.

What is the most reliable AI detector?

In independent benchmarks published from 2025 to 2026, Pangram has consistently ranked first, with Originality.ai also scoring well in the Authors Guild’s test of human writing. Re-test any tool on your own material, because accuracy changes as models evolve.

References