AI detectors are having their most public stress test yet. On the evening of Monday 21 September 2026, an anonymous account on X called Balance ton Claude posted screenshots claiming that Pangram, the best-regarded of the tools, had scored passages of C’était ça ou mourir as “100% AI”. The novel is the best-selling debut of Canadian-Haitian writer Thélyson Orélien, and it is in the running for the Prix Goncourt. Within a day the post had been viewed more than three million times, and a prize contender was on trial by software.
Orélien, 38, denies using AI. His French publisher, Grasset, calls the affair a hate campaign. And when AFP, franceinfo, Le Devoir and Radio-Canada ran the book through AI detectors of their own, the answers ranged from “very likely” human to 100% machine-written. That spread is the real story. It is the clearest natural experiment so far on a question that every teacher, editor, publisher and hiring manager now faces: how reliable are AI detectors, and what can a score actually prove?
This article works through the evidence. It sets out what each tester found in the Orélien case, how AI detectors work, what the independent benchmarks say, where the tools break down, and the arithmetic that decides whether a flag means anything. It builds on our earlier pieces on whether you can teach yourself to detect AI writing and on Pangram’s rise to gold-standard status. It ends with a practical policy for organisations that want to use AI detectors without wronging anyone.
Table of contents
- One Novel, Many Verdicts: The Orélien Case and AI Detectors
- How AI Detectors Actually Work
- Why AI Detectors Disagree on the Same Text
- What Independent Tests Say About AI Detectors
- The Arithmetic of False Positives
- Where AI Detectors Are Weakest
- The Human Cost of a Wrong Call
- How to Use AI Detectors Responsibly
- Beyond AI Detectors: Provenance, Watermarks and Disclosure
- Verdict: How Reliable Are AI Detectors in 2026?
- Frequently Asked Questions
- References
One Novel, Many Verdicts: The Orélien Case and AI Detectors
How the accusation started
C’était ça ou mourir (“It was that or die”) follows a Haitian migrant travelling to Canada. It is published by Grasset in France and by Boréal in Quebec. It won the Prix Fnac 2026, has been sold to publishers in more than 20 countries, and Le Devoir lists it among the finalists for the Goncourt, Médicis, Femina and Renaudot prizes this autumn. Doubts predated the X post. Podcaster Léa Bory suggested on Torchon that it read like AI because of repetitive sentences and a heavy use of analogies, and critic Fabrice Colin wrote in Le Canard enchaîné that he had “vainly looked for a unique voice”.
Balance ton Claude, whose name nods to Anthropic’s Claude model, had fewer than 100 followers before the post, according to franceinfo, and several thousand 24 hours later. “We tested several passages, at various points in the manuscript, for almost always the same result: 100% AI,” it wrote (our translation). It later said it had run four texts Orélien wrote in 2014 through Pangram and got 0% AI for each. It then urged people to go easy on the author, since there was a “non-zero chance” he was acting in good faith.
What the testers found with AI detectors
Four newsrooms and one poet then ran their own tests. The table below gathers every published result we could verify. None of the testers used identical text, tool settings or sample length, which is part of why the verdicts diverge so sharply.
| Tester | Tool | What was tested | Result |
|---|---|---|---|
| Balance ton Claude (X) | Pangram | Several excerpts, including the first two chapters | Almost always 100% AI |
| AFP | Several AI detectors | Excerpts of the book | From “very likely” human to 100% AI |
| franceinfo | Pangram, free version | Excerpts of varying length, dialogue and description, some mixing French with English and Spanish | 100% AI in every case |
| franceinfo | GPTZero | One excerpt | 87% chance of AI |
| Le Devoir | InVID-WeVerify, co-created by AFP | The passages cited on X | “Very likely” written by a human |
| Le Devoir | Seven other AI detectors | The same passages | One sided with Pangram, two said human, four leaned human without committing |
| Radio-Canada | Pangram | The full ebook | 94% of the text flagged as AI-generated |
| Poet Emmanuel Deraps | Pangram | One excerpt, then the same excerpt with about 20 words changed | 100% AI, then 100% human |
Read down the right-hand column and the problem is plain. Eight tests, at least a dozen AI detectors, and verdicts that span the whole scale. Some of that spread is noise from weak tools. Some of it is how the strongest tool is designed to behave. Separating the two is the job of the rest of this article.
The controls that complicate the story
The more careful testers also ran controls: texts whose authorship is not in doubt. Those results point the other way from the author’s defence, at least for Pangram on these samples.
| Tester | Control text run through Pangram | Result |
|---|---|---|
| Balance ton Claude | Four Orélien texts from 2014 | 0% AI each |
| franceinfo | Orélien’s 2015 book Le temps qui reste and a 2014 HuffPost article | 100% human |
| franceinfo | Extracts of four other autumn 2026 novels, by Ananda Devi, Philippe Jaenada, Amélie Nothomb and Lucie Rico | 99% to 100% human |
| Radio-Canada | Six other Goncourt-listed novels available digitally in Quebec | 100% human |
| Radio-Canada | Eight books by Haitian-born authors, including Lyonel Trouillot, Marie-Célie Agnant and Jacques Roumain, plus one by Blaise Ndala | 100% human |
| Radio-Canada | Whole chapters of the Bible and nine French classics, from Maupassant to Foucault | 100% human |
These controls answer the two defences raised most often. Orélien told Libération the software was “used to the French literary tradition and classical authors”. Grasset told BFMTV that detection tools “fabulate that the Bible was written artificially”. Marzena Karpinska of Simon Fraser University told Radio-Canada that some AI detectors do flag very short excerpts of much-quoted old texts, because those texts saturate the training data of language models. She said Pangram is not one of them, and that the novel still scored as AI when translated into English, which human texts translated the same way do not.
Radio-Canada later added a disclosure. Karpinska co-wrote a research paper with several Pangram employees, on AI-generated articles in US media, and Pangram donated credits for that work. She says she has no financial or professional tie to the company.
Where the case stands
None of this proves authorship, and Karpinska was careful to say so. “We cannot say with 100% certainty that it is AI,” she told Radio-Canada. “We can say there is a very high probability that this text was produced with significant AI assistance.” Entrepreneur Benoît Raphaël, who ran his own Pangram tests, told Nouvel Obs that a rating of “around 100% AI, or even 90%, is a pretty telling sign.”
On the other side, Orélien told La Presse that his first manuscript was finished in 2019, before ChatGPT existed. Boréal says the text matured through many rounds with his editor, “as the successive versions of the manuscript attest”. That version history is exactly the kind of evidence that could settle the question, and at the time of writing it had not been published. The sociologist Michel Lacroix told Le Devoir that a single nervous jury member could now sink the book’s prize chances, whatever the truth.
How AI Detectors Actually Work
Classifiers trained on paired text
Most commercial AI detectors are classifiers. They are trained on large collections of human writing alongside machine-generated text, and they learn the statistical features that separate the two. AFP’s summary is fair: the tools hold “large troves of AI-generated text alongside content written by humans” and compare the two.
Pangram’s twist, as Karpinska described it, is to pair the data. Its makers took millions of human-written texts, asked several AI models to rewrite the same material, and trained on those matched pairs. The model therefore learns how a human and a machine version of the same content differ, rather than learning topic or genre by accident. That design is one reason Pangram has pulled ahead of older AI detectors in independent tests.
Statistical scores and “linguistic tics”
Older and open-source tools lean on statistical scores such as perplexity: how predictable each next word looks to a language model. Machine text tends to be more predictable than human text, but so does plain, careful, conventional human prose. That overlap is the root of most false positives.
Thierry Poibeau, a researcher at France’s CNRS and a specialist in natural language processing, told AFP that the programmes look for “linguistic tics”. “The classic example is triadic phrasing (using three clauses), or em dashes,” he said. Those tells were common in early chatbot output because the models had been trained heavily on American scientific writing and American novels. Recurring words and phrases, and monotonous sentence structures, are further warning signs.
What the percentage on AI detectors means
A score such as “94% AI” is easy to misread. Radio-Canada’s 94% is the share of the book’s text that Pangram attributed to AI. It is not a 94% probability that the author cheated. GPTZero’s “87% chance” in franceinfo’s test is a different kind of number: a probability for one excerpt. Some AI detectors report a “human” percentage instead of an “AI” one, which is why the Authors Guild had to normalise every result before comparing tools.
The industry is slowly moving from binary verdicts towards estimates of how much AI went into a document, as we covered in our interview piece on Max Spero’s case for AI-assisted scoring. That is more honest, and harder to read at a glance. A reader who sees “94%” and thinks “94% guilty” has already made the most common mistake with AI detectors.
Why AI Detectors Disagree on the Same Text
Different tools, very different quality
The first reason is simply that AI detectors are not equally good. In the Authors Guild’s test, covered below, one consumer tool flagged every pre-AI human article as mostly machine-written. When a newsroom runs a passage through a dozen tools of mixed quality, the spread it reports is partly a spread in quality. Averaging a good detector with a bad one tells you very little.
Excerpts, length and free tiers
Short samples give a classifier less to go on. Turnitin raised its minimum from 150 to 300 words for exactly this reason, saying accuracy improves “with a little more text”. The University of Chicago benchmark found Pangram robust even on “stubs” of 50 words or fewer, but most AI detectors are not. The Orélien testers worked with excerpts of different lengths and, in franceinfo’s case, the free version of Pangram. Radio-Canada’s full-text run is the more informative of the tests.
Small edits, big swings
The poet Emmanuel Deraps told Le Devoir he replaced about ten words in a passage, deleted four or five and added four or five more. The first version scored 100% AI; the edited one scored 100% human. To many readers that looked like proof that AI detectors are random.
Karpinska’s reading is different. Light editing is how people “humanise” AI text to escape detection, she said, and it is easier to fool Pangram this way “because the tool is optimised to offer a minimal rate of false positives, while increasing the chance of false negatives”. In other words, a flip from AI to human after a light edit is the tool behaving as designed: it would rather miss AI text than accuse a human. The flip the other way, with human text scored as AI, is the rarer and more serious error.
Language, genre and literary style
Poibeau’s warning to AFP applies directly to fiction: “The tool may pick up something particular in the writing, something repetitive, but that might be exactly what constitutes a writer’s style.” Boréal made the same point in its statement, saying the tool’s reliability “has never, to our knowledge, been established on French-language novelistic prose”.
Karpinska told Radio-Canada that Pangram’s false-positive rate on French text is 0.0026%, and essentially zero on English literary text, but conceded there is no data specific to French literary prose. Her view is that fiction is easier, not harder: literary texts vary far more in form than news articles do, while AI tends to reproduce the same patterns. Filip Van Droogenbroeck of the Vrije Universiteit Brussel told Libération that Pangram deliberately trained on multilingual data, including under-represented languages.
Writers are not reassured. The French novelist Baptiste Beaulieu told Le Devoir he has always loved ternary rhythms and uses them everywhere. “So what do I do now?” he asked. “Change the way I write so that a piece of software agrees to believe I am human?”
What Independent Tests Say About AI Detectors
Vendor marketing is not evidence, so the useful question is what researchers with no stake in the outcome have found. The table summarises the six studies most often cited in this debate, from newest to oldest.
| Study | Tools tested | Sample | Headline finding |
|---|---|---|---|
| Vrije Universiteit Brussel, IJEI, June 2026 | GPTZero, Pangram, Copyleaks, Turnitin | 160 synthetic academic papers; 1,163 master’s theses | Pangram best; the others underestimated AI content; all four identified fully human texts |
| Authors Guild, May 2026 | Pangram, Originality.ai, Grammarly, ZeroGPT, Sidekicker | Ten Guild articles from 2022 or earlier | Pangram 0% AI on all ten; Sidekicker 71% to 100% AI on all ten |
| University of Chicago, NBER w34223, 2025 | Pangram, Originality.ai, GPTZero, RoBERTa | 1,992 human and 1,992 AI texts across genres, lengths and models | Pangram near-zero error rates; the only tool to meet a 0.5% false-positive cap |
| Russell, Karpinska and Iyyer, ACL 2025 | Pangram, GPTZero, Fast-DetectGPT, Binoculars, RADAR, human experts | 300 articles | Pangram caught 98.0% of AI articles; RADAR caught 15.3% |
| Liang et al., Stanford, Patterns, July 2023 | Seven early AI detectors | 91 TOEFL essays plus US eighth-grade essays | 61% of non-native essays falsely flagged, on average |
| OpenAI, January to July 2023 | OpenAI’s own classifier | OpenAI’s challenge set | Caught 26% of AI text and falsely flagged 9% of human text; withdrawn |
The University of Chicago benchmark
The paper AFP cites, “Artificial Writing and Automated Detection” by Brian Jabarian and Alex Imas, is the most rigorous head-to-head of commercial AI detectors so far. It tested three commercial tools and one open-source model across a large corpus spanning genres, lengths and generating models. The commercial tools beat the open-source one, and Pangram achieved “near-zero” false-negative and false-positive rates that held across models, thresholds, ultra-short passages and “humanizer” tools.
Its most useful contribution is a framework rather than a ranking. The authors let a decision-maker set a “policy cap”, a maximum tolerable error rate that does not depend on the tool. Under a strict cap of 0.5% false positives, Pangram was the only one of the four AI detectors that met it without sacrificing accuracy.
The Vrije Universiteit Brussel study
Published in the International Journal for Educational Integrity in June 2026, the VUB study built 160 academic papers with known ground truth: fully human, fully AI, hybrid, and AI text “humanised” by a prompt designed to resemble what a student might do. Pangram consistently outperformed GPTZero, Copyleaks and Turnitin. The other three “significantly underestimated” AI content, especially from the newest model.
The authors then ran Pangram over 1,163 real master’s theses from the 2024-2025 academic year. It flagged 45.5% of them, typically at low to moderate levels of AI-associated text. False positives were rare across all tools in the controlled set, yet the authors still concluded that AI detectors “should not be used as sole evidence in high-stakes decision-making”.
The Authors Guild test of AI detectors
In May 2026 the Authors Guild, the largest professional body for US writers, ran ten of its own articles, all published in 2022 or earlier, through five tools. Any accurate tool should have scored every one as human. The chart shows the highest AI score each tool gave to any of the ten.
The detail is worse than the maximums. ZeroGPT scored the Guild’s obituary of Joan Didion at 66% AI and its congratulations to Louise Erdrich on her Pulitzer at 76%, with no pattern the Guild could explain. Sidekicker scored every article between 71% and 100%. The Guild’s verdict: “some commonly used consumer-facing AI detection tools are wildly inaccurate”, and even the best “should never be the sole basis for any decision”.
The expert-reader study
Jenna Russell, Marzena Karpinska and Mohit Iyyer tested five AI detectors against 300 articles at ACL 2025, and compared them with human readers. At a fixed 2% false-positive rate, Pangram caught 98.0% of the AI-written articles, GPTZero 85.3%, Fast-DetectGPT 80.0%, Binoculars 66.7% and RADAR 15.3%. A panel of frequent chatbot users, voting by majority, missed just one AI article in 300. This is the “2024 test” Karpinska described to Radio-Canada, and it is where Pangram’s reputation began.
The older warnings
Two 2023 results still shape the public view. OpenAI launched its own classifier on 31 January 2023, reported that it caught only 26% of AI-written text while falsely flagging 9% of human text, and withdrew it on 20 July 2023 “due to its low rate of accuracy”. That same year, Weixin Liang and colleagues at Stanford found seven popular AI detectors wrongly flagged 61% of TOEFL essays by non-native English speakers on average, while passing more than 90% of essays by US eighth-graders.
Put together, the pattern is clear. Since 2023 the field has split. A small number of AI detectors, with Pangram the most consistent, now post very low false-positive rates on long English text in independent tests. Many widely used consumer tools do not. So “are AI detectors reliable?” has no single answer. It depends on which tool, which text and which error you care about.
The Arithmetic of False Positives
Two errors with two different harms
The Chicago paper frames the problem in two numbers. The false-negative rate is the share of AI-generated text a tool wrongly passes as human. The false-positive rate is the share of human-written text it wrongly flags as AI. A school worried about cheating cares about the first. A novelist accused of cheating cares about the second. Most AI detectors let the operator trade one against the other by moving a threshold, and Pangram, by Karpinska’s account, is deliberately tuned to keep false positives minimal at the cost of missing more AI text.
Vendor claims, and why they are hard to check
AFP made a blunt observation: many tools publish percentage figures about their accuracy, but the figures are impossible to verify and do not account for the risk of a false positive. The published numbers that do exist come from very different tests.
Turnitin, after running 800,000 pre-ChatGPT academic papers through its system, said its document-level false-positive rate is under 1% for documents where it detects more than 20% AI writing, and that its sentence-level rate is about 4%. Scores below 20% now carry an asterisk because they are less reliable. Pangram’s own site put its overall rate at about 1 in 10,000 (0.01%) as of 15 September, and 0.003% for books, according to Radio-Canada. Earlier this year the company reported 0.0041% for its Pangram 4 model. Vendor figures move, and they are measured on the vendor’s own test sets.
To make those rates concrete, multiply them out. The chart shows how many human-written documents out of 10,000 each stated false-positive rate would wrongly flag.
The gap between the best and worst AI detectors is three orders of magnitude. At Turnitin’s upper bound, a university that screens 10,000 honest essays could still produce up to 100 false flags. Behind each, as Turnitin itself wrote, “is a real student who may have put real effort into their original work”.
Base rates decide what a flag means
A false-positive rate alone does not tell you how likely a flagged document is to be innocent. That depends on how common AI writing is in the pile you are screening. The worked example below takes a hypothetical detector that catches 90% of AI text and falsely flags 1% of human text, Turnitin’s stated upper bound, and applies it to 10,000 documents at three levels of AI use. The last row repeats the lowest level with a 0.01% false-positive rate.
| Share of AI-written documents | AI documents caught | Human documents flagged | Flags that are wrong |
|---|---|---|---|
| 2% (200 of 10,000) | 180 | 98 | 35.3% |
| 10% (1,000 of 10,000) | 900 | 90 | 9.1% |
| 30% (3,000 of 10,000) | 2,700 | 70 | 2.5% |
| 2%, with a 0.01% false-positive rate | 180 | about 1 | 0.5% |
The lesson is uncomfortable. Where AI use is rare, as it should be among shortlisted literary novels, even a 1% false-positive rate means more than a third of flags point at innocent writers. Only a far lower rate makes a flag meaningful in a low-prevalence pile. That is why the choice of tool matters so much, and why the same score means different things in a first-year essay queue and a prize shortlist.
Where AI Detectors Are Weakest
Short texts, openings and endings
Every serious study agrees that length helps. Turnitin found more false positives in the first and last few sentences of a document, typically the introduction and conclusion, and changed how it aggregates them. OpenAI said its classifier was “very unreliable” below 1,000 characters. Excerpts pasted into a free web form, the way most people test a suspicion, are the least reliable input AI detectors see.
Humanisers and light edits
An entire industry sells rewriting tools that exist to beat AI detectors. We have reviewed two of them, Go Humanize AI and Rehumanize.io. The Chicago benchmark found Pangram robust to the humanisers it tested, while the VUB study found the other three tools badly underestimated humanised text. The Deraps experiment shows the limit: a person making careful edits by hand can push a passage below the threshold of a tool tuned to avoid false positives.
Non-native and highly polished writers
The Stanford result still defines the bias debate. Early AI detectors punished the simpler, more predictable phrasing typical of people writing in a second language. The Authors Guild describes a mirror-image problem at the other end of the skill scale: every large language model was trained on polished, edited prose, so “the more refined and controlled a writer’s style, the more it may resemble the output these tools are designed to flag”. Neurodivergent students and writers with a strongly patterned voice raise the same concern.
The moving target
Poibeau calls it the “moving target problem”: AI programmes are “constantly evolving”, so a tool trained on last year’s models may miss this year’s. The Authors Guild makes the same point about the detectors themselves, whose accuracy “at any given moment cannot be assumed”. A result from 2023, good or bad, says little about a tool in 2026, and a result from today may not hold after the next model release.
| Weak spot | Main error it causes | Evidence |
|---|---|---|
| Short excerpts | Both errors rise | Turnitin raised its minimum to 300 words; OpenAI warned below 1,000 characters |
| Light human edits | False negatives | About 20 changed words flipped a Pangram result from 100% AI to 100% human |
| Humaniser tools | False negatives | VUB: three of four tools underestimated humanised text |
| Non-native writers | False positives | Stanford: 61% of TOEFL essays flagged by early tools |
| New models | False negatives | VUB: tools other than Pangram struggled most with the newest model |
The Human Cost of a Wrong Call
For authors and publishers
Orélien is not the first writer caught in this. In March, Hachette withdrew the horror novel Shy Girl days before its planned US release after Pangram identified the text as AI-written, Radio-Canada reports. In May, one of the five winning stories of the 2026 Commonwealth Short Story Prize, “The Serpent in the Grove” by Jamir Nazir, scored 100% AI on Pangram. The Commonwealth Foundation and Granta stood by it, noting that AI detectors are imperfect, while the Foundation reviews its selection process.
The Authors Guild’s advice to publishers is firm. They should disclose their methods, keep up with the tools’ accuracy, and give authors “a fair and full opportunity to defend themselves”. They “should never cancel a contract or pull a book on the basis of accusations or use of these tools alone”, which the Guild says would be a material breach of contract.
For students and teachers
Education is where AI detectors are used most. By May 2023, Turnitin had already processed 38.5 million submissions for AI writing. The VUB theses result shows what that looks like in practice: nearly half of real master’s theses carried some flag, mostly at low levels. Each of those flags is a conversation a supervisor has to handle well, and a student who may be innocent.
For businesses, hiring and reviews
The Chicago authors open their paper with business uses: checking that product reviews were written by real customers, not only that coursework was done by students. The same logic now reaches cover letters, supplier proposals, commissioned copy and customer complaints. Pangram’s chief executive, Max Spero, told The Atlantic that the tool should not be the ultimate arbiter, but a starting point for a more in-depth investigation. Organisations using AI detectors for screening should take the vendor at its word.
How to Use AI Detectors Responsibly
The evidence supports a narrow, useful role for AI detectors: triage. The table sets out what a detector score can reasonably support in the situations where organisations meet it most.
| Situation | Reasonable use of a score | Unreasonable use |
|---|---|---|
| Large inbound queue (reviews, applications) | Prioritise items for human review | Reject flagged items automatically |
| Academic submission | Open a conversation and ask for drafts | Sole basis for a misconduct finding |
| Manuscript or commissioned copy | Request version history and process evidence | Cancel a contract or pull a book |
| Text under 300 words | Treat as a weak signal at most | Any decision at all |
| Second-language or highly polished writer | Weigh against known bias and context | Treat a high score as proof |
Treat a score as a lead, not a verdict
Every independent study reviewed here, and the leading vendor itself, says the same thing: a score opens an inquiry and never closes one. Write that into the process, so that no adverse decision can be taken on a detector result without a named person reviewing the wider evidence.
Choose AI detectors on independent evidence
Pick a tool on third-party benchmarks, not the accuracy figure on its landing page. On today’s evidence, that means a tool with a published, independently tested false-positive rate on text like yours. Then test it on your own material. Run a sample of documents you know are human, such as work written before 2022, and a sample you know is AI-generated. If it fails your own controls, its marketing figures are irrelevant.
Test the whole document and run controls
Radio-Canada’s approach is the model. It tested the full text rather than excerpts, then ran the same tool over comparable books whose authorship is not in doubt. When a result matters, do both. A high score on a full document, set against clean scores on the same writer’s earlier work and on comparable texts, is far stronger evidence than one pasted paragraph.
Ask for process evidence
The best evidence of authorship is usually not in the text at all. Version history in a word processor, dated drafts, notes, research trails and an editor’s correspondence all show how a piece came to exist. Ask for them early, and ask neutrally. Orélien’s publishers say such a record exists for his novel, and producing it would say more than any of the AI detectors have.
Write a detection policy
Decide in advance which uses of AI are allowed, which tool you use, what threshold triggers review, who reviews, and how people can respond. Our AI governance framework for SMEs covers the wider policy structure, and our AI strategy team can help you set one up. A written policy protects the organisation as well as the people it screens.
Beyond AI Detectors: Provenance, Watermarks and Disclosure
Watermarks at the point of generation
AI detectors guess after the fact. Watermarking marks text as it is generated, by nudging a model’s word choices into a pattern that can be tested for later. Google has open-sourced its SynthID text method, as we explained in our guide to verifying AI content with SynthID. Watermarks only work for models that apply them, and paraphrasing weakens them, but a positive result is far stronger evidence than a statistical guess.
Labelling duties
Regulators are moving the burden onto providers. Since 2 August 2026, Article 50 of the EU AI Act has required providers to mark synthetic output in a machine-readable way, and deployers to disclose deepfakes and some AI-generated text. We covered the details, and the risks of labelling, in our piece on the EU’s AI content label rules. None of this helps with a novel written by a person using a chatbot privately, which is exactly the Orélien scenario.
Disclosure as a norm
Even Balance ton Claude ended on disclosure. It said it flagged the novel because it was successful, that it was “far from the only one”, and called on authors who use AI to say so rather than betray their readers. The Authors Guild already runs a Human Authored certification for books, and publisher contracts increasingly ask about AI use. They do not replace AI detectors, but they change the question from “can we catch it?” to “did you tell us?”
Verdict: How Reliable Are AI Detectors in 2026?
What the evidence supports
The best AI detectors are now far better than their reputation. On long English text, independent studies in the US and Belgium found Pangram’s false-positive rate close to zero, and several found it reliably catches unedited AI output. In the Orélien case its result was consistent across excerpts, the full book and translation, while its controls on comparable human books came back clean.
What it does not support
The typical free tool is still unreliable, and even the best can be pushed under its threshold by careful editing. No study has established error rates for French literary prose. Scores are routinely misread, excerpts are routinely used in place of full texts, and the arithmetic of base rates means a rare false positive still lands on real people. A score, however high, is not proof of how a text was written.
The bottom line
So, how reliable are AI detectors? The good ones are reliable enough to justify a closer look, and not reliable enough to justify a verdict on their own. That is the same answer researchers, the Authors Guild and Pangram’s own chief executive give. The Orélien affair will be settled, if at all, by drafts and version history, not by another screenshot.
Frequently Asked Questions
Are AI detectors accurate?
It depends heavily on the tool. Independent tests from the University of Chicago and the Vrije Universiteit Brussel found Pangram’s error rates near zero on long English text. The Authors Guild found some consumer tools flagging every human article it tested as mostly AI, with scores up to 100%.
Why did AI detectors give different results for the same novel?
The testers used different tools, different excerpts and different versions. Tool quality varies widely, short excerpts are less reliable, and scores measure different things: a share of text, a probability or a “human” percentage. A light edit of about 20 words also flipped one Pangram result.
Can AI detectors be fooled?
Yes. Humaniser tools and careful hand edits can push AI text below a detector’s threshold, especially for tools tuned to minimise false positives. Evading detection is easier than being falsely accused by the best tools, though weaker tools do falsely accuse human writers.
Do AI detectors discriminate against non-native English speakers?
Early ones did. A 2023 Stanford study found seven popular tools falsely flagged 61% of TOEFL essays on average. Newer tools report much lower rates, but the Authors Guild warns that highly polished human prose can also resemble AI output.
Should a publisher, school or employer act on a detector score alone?
No. Independent researchers, the Authors Guild and Pangram’s chief executive all say a score should start an investigation, not end one. Ask for drafts and version history, test full documents, and run controls before reaching any conclusion.
What is the most reliable AI detector?
In independent benchmarks published from 2025 to 2026, Pangram has consistently ranked first, with Originality.ai also scoring well in the Authors Guild’s test of human writing. Re-test any tool on your own material, because accuracy changes as models evolve.
References
Spotting AI writing: how reliable are the detectors? (AFP via IOL)
Controverse Thélyson Orélien : ce qu’il faut savoir sur Pangram (Radio-Canada)
Thélyson Orélien nie avoir utilisé l’IA pour son roman (Radio-Canada)
Thélyson Orélien accusé sur X d’avoir généré son roman par IA (Le Devoir)
Vrai ou faux : Thélyson Orélien a-t-il utilisé l’IA pour écrire son premier roman ? (franceinfo)
Artificial Writing and Automated Detection (NBER Working Paper 34223)
Can AI Detectors Be Trusted? The Authors Guild Put Five of Them to the Test
GPT detectors are biased against non-native English writers (Liang et al.)
New AI classifier for indicating AI-written text (OpenAI)
AI writing detection update from Turnitin’s Chief Product Officer
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.