Detect AI writing by eye — no tool, no upload, just your own reading — and you are doing something that most researchers assumed, until quite recently, could not be done at all. The assumption was reasonable. Study after study put ordinary readers at roughly coin-flip accuracy, and the machines built for the job kept flagging innocent people. Then a set of experiments handed the same task to people who use these systems every day, and the numbers moved sharply.

That is the honest shape of the answer, and it is the one Celeste Rodriguez Louro, Associate Professor and Director of the Language Lab at the University of Western Australia, arrived at in a piece published in The Conversation on 25 August 2026. Can you teach yourself to spot synthetic prose? Maybe. Not through one magic giveaway, but through the slow accumulation of habits — and only if you accept that every habit you learn has an expiry date. Models are tuned constantly, sometimes with reinforcement learning against human preference, and a tell that works this month can be trained away by the next release.

The stakes are no longer academic. A debut novelist lost a seven-figure deal this summer because nobody could prove how his manuscript came together. Teachers are marking work they cannot vouch for. Editors are commissioning writers they cannot verify. And the tools sold to settle these arguments have a documented bias problem that lands hardest on people writing in a second language. If you track this space through our AI models, tools and releases hub, this is the part of the story where reading skill, rather than software, turns out to carry the weight.

Below: the vocabulary and structural tells that genuinely correlate with machine output, the accuracy figures for expert readers against commercial detectors, the bias data that should stop you accusing anyone on a hunch, and a practical routine for training your own eye without deceiving yourself about how good it is.

Why Detect AI Writing Is Suddenly A Business Problem

detect ai writing tells and limits b sieve bowl on plinth

For two years this was a staffroom argument. In 2026 it started ending careers and contracts, which is a very different thing.

A seven-figure deal withdrawn on suspicion alone

Nigerian author Jerry Falade’s debut crime novel, Call Me, I’ll Hide the Body, drew a 14-way auction in the United States, with Minotaur — a Macmillan US imprint — reportedly offering more than $2 million. According to The Guardian on 31 July 2026, his own agents then pulled the manuscript from submission, saying they could no longer authenticate how it had evolved. The agency cut ties after a meeting on 29 July.

Nobody produced proof either way

Falade denies using AI in the writing or editing, and has argued that the reaction to his work was shaped by racial bias. An editor who bid on the book told The Bookseller they detected no AI while reading it, and neither did the sales, marketing and publicity colleagues who read alongside them. No detector output settled it. The deal collapsed on doubt, not evidence.

It was not an isolated case

In March 2026 Hachette withdrew the horror novel Shy Girl under similar suspicion. Two withdrawals in one publishing year, neither resolved by a test, tells you something important: the industry currently has no accepted method to verify authorship. That vacuum is exactly why people are trying to learn to detect AI writing themselves.

What It Costs To Detect AI Writing Wrongly

Miss a synthetic manuscript and you publish something your contract says you did not. Wrongly accuse a human writer and you damage a career on a hunch. Both failure modes are live, and only one of them currently generates headlines. Any policy that asks staff to detect AI writing has to say what happens next when somebody thinks they have.

The Vocabulary Tells That Help You Detect AI Writing

detect ai writing tells and limits c mirror upright stand

There is no single giveaway. There are habits, and the strongest ones are lexical — particular words that these systems reach for far more often than people do.

TellWhat it looks likeHow reliable on its own
Signature vocabularydelve, meticulous, underscore, boast, intricateModerate, and fading
Quiet adverbsquietly, genuinely, significantlyModerate
Not X, but Y“not a feature, but a philosophy”Moderate
The rule of threeThree-item lists, everywhereWeak alone
Claim escalationModest point, immediate grand conclusionModerate
Length mismatchPolished paragraphs where a line would doStrong in context
Flawless mechanicsNo typos, no ragged edges, everWeak alone
Em dashesHeavy use in one model, none in anotherPoor

The words that spiked in scientific abstracts

Researchers tracking PubMed abstracts have documented sharp rises since 2022 in delve, meticulous, underscore, boast and intricate. Nothing is wrong with any of them. What is unusual is the rate — a vocabulary shift across an entire corpus, arriving in step with the tools. When you try to detect AI writing in formal prose, that cluster is the first thing worth noticing.

Quietly, genuinely, and the escalation habit

Rodriguez Louro singles out quietly as a strong signal; search interest in the word has climbed on Google Trends since December 2021. Genuinely shows the same inflation. Then there is the rhetorical move underneath the vocabulary: take a modest claim and immediately escalate it into something sweeping. Human writers do this occasionally. Machines do it with remarkable consistency, paragraph after paragraph.

Why the em dash never really worked

The em dash was an early favourite tell and it never held up. One comparison found ChatGPT producing eight em dashes in 573 words while Gemini and Meta AI produced none in the same task. A marker that varies that much between systems cannot help you detect AI writing on its own — and plenty of human writers, this one included, have used the em dash freely for decades.

Why Vocabulary Alone Cannot Detect AI Writing

Every item above is a base-rate observation, not a proof. A careful academic who genuinely likes the word underscore is not a chatbot. The value is cumulative: three or four of these habits stacked in one piece is a signal worth acting on; one of them is noise.

Structural Habits That Help You Detect AI Writing

detect ai writing tells and limits d upright thermometer

Expert readers in the research do not stop at vocabulary. They read for shape, and shape is harder for a model to disguise than word choice.

Sentence architecture that repeats

The not only… but also frame, the tricolon, the tidy topic-sentence-then-three-supports paragraph. Machine prose is built from a small set of reliable moves, deployed relentlessly. Human writing wanders, backtracks, and occasionally builds a sentence that does not quite work. That unevenness is information, and readers who detect AI writing well tend to feel its absence before they can name it.

Length that ignores the situation

Ask a person how the meeting went and you get a fragment. Ask a model and you get four balanced paragraphs. Rodriguez Louro’s clearest everyday example is the text message that arrives in polished prose where shorthand belongs. Register mismatch — formal machinery applied to an informal moment — is one of the most practical ways to detect AI writing in ordinary correspondence.

Quotations that sound wrong

Annotators in the research repeatedly flagged quotes: sources who all speak in the same measured register, with none of the hesitation, slang or specificity of real speech. A quote that reads like a press release inside an article that reads like a press release is a compounding signal.

Grammar that is too clean

Machine output is usually grammatically perfect. That is a weak tell alone, because plenty of professional copy is clean too. It becomes useful in combination — perfect mechanics plus safe, surprise-free content plus signature vocabulary is a very different pattern from perfect mechanics on its own. Stacking is how you detect AI writing without leaning on any one clue.

Originality, or the absence of it

The category experts named most after vocabulary and structure was originality: writing that is straightforward, safe and free of anything unexpected. Nothing in it that a competent generalist would not have said. This is subjective, and it is also, in the data, one of the things human readers assess far better than software — which is much of the reason people can detect AI writing that a detector waves through.

How Well People Actually Detect AI Writing

detect ai writing tells and limits e tuning fork upright

Here the evidence gets specific, and considerably more encouraging than the old consensus.

Jenna Russell, Marzena Karpinska and Mohit Iyyer hired annotators to read 300 non-fiction English articles — a mix of human-written pieces and output from GPT-4o, Claude and o1 — label each one, and explain their reasoning in a paragraph. The study ran 1,790 annotations at a total cost of about $4,900, and was published at ACL in 2025.

The result that matters: annotators who frequently use these systems for their own writing were excellent at the task, with no specialised training at all. Below are the five expert annotators’ true-positive rates on the first experiment, against GPT-4o output.

Expert annotators, true-positive rate on GPT-4o articles (Russell et al., 2025)
Annotator 1 96.7%
Annotator 2 96.7%
Annotator 5 93.3%
Annotator 4 90.0%
Annotator 3 86.7%
Non-expert average 56.7%

Five readers voting together missed one article in 300

Pooled as a majority vote, the five experts reached 99.3% true-positive rate at a 0% false-positive rate across the experiments — one misclassification out of 300 articles. That is not coin-flip performance. A panel that can detect AI writing at that rate beats most of the software evaluated in the same study, including under evasion tactics such as paraphrasing.

Occasional Users Could Not Detect AI Writing Reliably

The same task given to annotators who rarely use these systems produced a 56.7% true-positive rate against a 51.7% false-positive rate. In other words: they said “AI” about as often for human text as for machine text. The difference between the two groups is not intelligence or education. It is exposure — how much machine text they had already read. Occasional users could not detect AI writing at any rate worth acting on.

Exposure Is How People Learn To Detect AI Writing

That is the encouraging finding buried in the numbers. The experts had no training programme, no rubric and no feedback loop. They simply read a great deal of machine output in the course of their own work. If you want to detect AI writing well, the evidence says the fastest route is volume — read a lot of it, deliberately, knowing what it is.

Where the experts named their reasoning

Their written explanations were coded into categories, and the ranking doubles as a checklist for anyone learning to detect AI writing.

Clue categoryShare of expert explanationsTypical wording
Vocabulary53.1%Specific words used far more than people use them
Sentence structure35.9%Predictable frames such as “not only… but also”
Grammar and punctuation24.8%Grammatically perfect, without exception
Originality23.7%Safe, straightforward, lacking surprises
Quotations22.3%Overly formal, short on nuance

Where Software Detects AI Writing Better Than People

detect ai writing tells and limits f two identical blank cards

The same study benchmarked commercial and open-source detectors on the identical articles, which makes it one of the few genuinely like-for-like comparisons available.

True-positive rate, human panel against automated detectors (Russell et al., 2025)
Expert majority vote 99.3%
Pangram 98.0%
GPTZero 85.3%
Fast-DetectGPT 80.0%
Binoculars 66.7%
RADAR 15.3%

One tool kept pace with the panel

Pangram reached 98.0% true positives at a 2% false-positive rate, and held 99.3% on the humanised set. It is a genuine outlier, and the only tool in the set that could detect AI writing about as well as the panel. The rest of the field trailed: GPTZero at 85.3%, Fast-DetectGPT at 80.0%, Binoculars at 66.7%, and RADAR at 15.3% — a tool that, on this corpus, missed roughly six machine articles in every seven.

False positives are the number to read

True-positive rate is the flattering half of the picture. The false-positive rate is where damage lives, because that is the innocent writer being accused. The expert panel’s 0% on that measure is the single most striking figure in the study. People who detect AI writing carefully accused nobody wrongly; every automated system tested did, including the one that matched them on detection.

Humanisers break some readers too

Run machine text through a paraphrasing “humaniser” and results scatter. One expert annotator’s true-positive rate collapsed from 16.7% to 0% on the humanised o1-Pro set, while another held at 100%. The panel survived because majority voting absorbed individual failures. A single reader trying to detect AI writing alone, on one document, has no such safety net.

The Bias That Should Stop You Accusing Anyone

If you take one thing from the research, take this. The tools sold to settle authorship disputes have a documented, measured, unresolved fairness problem.

Group testedWhat the detectors did
91 TOEFL essays, non-native writers61% falsely flagged as AI on average, across seven tools
Worst single tool, same essaysClose to 98% falsely flagged
US eighth-grade essays, native writersOver 90% correctly classed as human

The Stanford finding, in plain terms

Weixin Liang and colleagues ran seven popular detectors over 91 essays written by non-native English speakers sitting the TOEFL exam. On average, 61% were wrongly labelled machine-generated; one tool flagged nearly 98% of them. The same detectors classified more than 90% of essays by US eighth-graders correctly as human. Published in Patterns in July 2023, it remains the clearest warning against outsourcing the job of detecting AI writing to software.

Why simpler prose reads as synthetic

The mechanism is not mysterious. These detectors lean on statistical measures of surprise — how predictable the next word is. Writing with a smaller vocabulary and more conventional phrasing scores as predictable, and predictability is what the software calls machine-like. Second-language writers are penalised for exactly the features that make their English harder-won, which is why software that claims to detect AI writing fails them first.

What To Do When You Detect AI Writing In Someone’s Work

A detector score is not evidence. Neither is your own reading, however practised. Falade’s case is the cautionary version: a career-altering decision made on unprovable suspicion, in a field with no verification standard, against a writer who has attributed the reaction to bias. Whatever you conclude when you detect AI writing in someone’s work, treat it as a question to ask, never a finding to announce.

Why The Tells You Use To Detect AI Writing Keep Expiring

Every marker described above is a snapshot of a moving target, and the target moves faster than folk knowledge spreads.

The Words That Detect AI Writing Get Trained Away

These systems are updated rapidly and cheaply. When delve became a public joke, it became a thing to tune away. Anything widely discussed as a giveaway has a short half-life precisely because it is widely discussed. The vocabulary list you memorise this quarter will detect AI writing from last quarter’s models, and not much else.

The homogenisation problem cuts the other way

There is a deeper trend documented in Nature Human Behaviour. Zhivar Sourati and colleagues analysed more than 880,000 texts — Reddit stories, news articles, academic papers, essays, social posts and political speeches — rewritten by GPT-3.5, Llama 3 70B and Gemini Pro. Rewriting cut variation in writing complexity by 21% to 50%, while 87% of pairs kept meaning-similarity scores above 0.95.

MeasureFinding
Corpus sizeMore than 880,000 texts across six genres
Variation in writing complexityReduced by 21% to 50% after rewriting
Meaning preserved87% of pairs scored above 0.95 similarity
Author-trait prediction6 percentage points less accurate afterwards

Personal signal is being sanded off

After rewriting, models trying to infer authors’ personal characteristics from their language were on average six percentage points less accurate. Associations that normally hold — pronoun use with extraversion, friend-related words with loyalty, future-focused words with age — were weakened. The individuality that lets a careful reader recognise a voice is exactly what the rewriting removes — and voice is a large part of what lets anyone detect AI writing at all.

Provenance is where the regulation is heading

The European Union now requires AI-generated content to be labelled or watermarked. That is a bet that provenance will scale where forensics will not — and it is the right bet, because a label travels with the file while a tell has to be re-learned every release cycle. Learning to detect AI writing by reading remains useful, but it is a stopgap, not an infrastructure.

How To Train Yourself To Detect AI Writing

The research supports a specific and slightly unglamorous method. It is mostly reading.

Read Enough Output To Detect AI Writing By Feel

The experts in the study had one thing in common: heavy daily use. Generate text on topics you know well, in registers you write in, and read it closely. You are building a base rate for what median output feels like, and that base rate — not a checklist — is what lets you detect AI writing quickly.

Score on accumulation, never on one clue

Treat each marker as a small weight rather than a switch. One delve is nothing. Signature vocabulary plus escalating claims plus flawless mechanics plus a length that ignores the context is a genuine pattern. This is how the expert annotators worked, and it is why they beat tools that keyed on narrower features. Nobody who can detect AI writing reliably is doing it from one word.

Write down your reasoning before you check

Force yourself to produce the paragraph of explanation the annotators were required to write. Articulating why converts a vague feeling into a testable claim, and it exposes the times your only real evidence was that the prose was tidy. Then check, and keep score honestly.

Calibrate on writing that is not yours

Test yourself on second-language English, on technical documentation, on corporate copy — the registers most likely to trip you into a false positive. If you only detect AI writing accurately in the genres you already know, you have learned a genre, not a skill.

Accept the ceiling

Even the panel that missed one article in 300 was five people voting, on long-form non-fiction, against known models. Alone, on a short document, from an unknown model, you will be worse than that. Knowing roughly how much worse you detect AI writing under those conditions is the difference between a useful reader and a dangerous one.

What This Means If You Publish Or Commission Content

For a business, the practical answer is to stop treating this as a forensics problem and start treating it as a process problem.

Provenance Beats Any Attempt To Detect AI Writing

Version history, drafts, research notes, commit logs, dated outlines — the things Falade’s agents said they could not reconstruct. A writer who can show the road from brief to final copy does not need to pass a test. Ask for that at commissioning, not during a dispute. It is cheap in advance and impossible afterwards.

Write policy about disclosure, not about tools

A policy that says “we run everything through a tool that claims to detect AI writing” inherits that tool’s bias and its false-positive rate. A policy that says what assistance is permitted, and requires it to be declared, is enforceable and fair. Our AI strategy work with clients starts at this distinction, because getting it backwards creates disputes rather than settling them.

Editorial judgement is still the control

Homogenised, surprise-free prose is a quality problem before it is an authenticity problem — and it is a ranking problem too, since search and answer engines increasingly reward distinctive, first-hand material. Editing for specificity, first-hand detail and an actual point of view fixes all three at once, and it matters more than being able to detect AI writing after the fact. That is the same standard our SEO services and AEO services teams apply to client content, whatever produced the first draft.

Train the people who read, not just the ones who write

If exposure is the intervention that works, then the editors and managers reviewing copy should be using these tools themselves. The organisations whose reviewers can detect AI writing reliably will be the ones whose reviewers use the technology daily — a slightly awkward conclusion, and a well-evidenced one.

Your Next Steps

Start with the two-week version of the experiment. Generate a few hundred words a day on subjects you know well, read them properly, and note what recurs. Then take real documents you already know the origin of and score yourself, writing your reasoning down first. Most people find their accuracy is high on long-form prose and poor on anything short.

Then fix the process while you are learning the skill. Ask new writers for drafts and version history as a matter of course. Put a disclosure clause in your contributor terms. And if you find yourself about to accuse someone on the strength of a detector score or a hunch about the word quietly, reread the 61% figure first. The evidence says you can teach yourself to detect AI writing, up to a point. It says nothing that justifies certainty.

Frequently Asked Questions

Can you really learn to detect AI writing by reading?

Yes, partially. Annotators who used these systems heavily reached 86.7% to 96.7% true-positive rates individually and 99.3% as a majority vote, with no training programme. Occasional users scored 56.7%, barely above chance. Exposure to machine text, not instruction, is what let one group detect AI writing and not the other.

What is the single most reliable tell?

There isn’t one, and hunting for a single tell is the main mistake people make when they try to detect AI writing. Vocabulary was cited in 53.1% of expert explanations, which makes it the strongest single family of clues, but the experts always reasoned from several signals at once.

Are AI detectors accurate enough to rely on?

Mostly no. In one head-to-head, Pangram reached 98.0% true positives, but GPTZero managed 85.3%, Fast-DetectGPT 80.0%, Binoculars 66.7% and RADAR 15.3%. Separately, seven detectors falsely flagged 61% of essays by non-native English writers.

Is the em dash a sign of AI writing?

Not a dependable one. In one test ChatGPT used eight em dashes in 573 words while Gemini and Meta AI used none. It varies too much between systems, and many human writers use it heavily.

Why do detectors flag second-language writers so often?

They score how predictable the wording is, and more conventional phrasing looks predictable. Liang and colleagues found 61% of TOEFL essays by non-native speakers wrongly flagged, one tool reaching almost 98%, while over 90% of native eighth-grade essays passed.

Will these tells still work next year?

Some will not. Models are updated quickly, and any marker discussed publicly becomes a candidate for tuning away. Treat the specific words you use to detect AI writing as perishable, and the underlying habits — escalation, register mismatch, absence of surprise — as more durable.

References