Publishing

ai detectors spotting ai writing how reliable a searchlight on a yoke and pedestal

Spotting AI Writing: How Reliable Are the Detectors?

When an X account used Pangram to accuse Canadian-Haitian novelist Thelyson Orelien of writing his Goncourt-listed debut with AI, newsrooms ran the book through a dozen AI detectors and got every possible answer. This article sets out what each test found, how the tools work, what independent benchmarks from Chicago, Brussels and the Authors Guild show, the arithmetic of false positives, and how to use AI detectors without wronging anyone.

Read more
pangram ai detection gold standard trust a solid rubber stamp thick handle

Pangram Has Emerged as the Gold Standard of AI Detection. Should You Trust It?

Pangram, a 24-person Brooklyn startup, has become publishing’s de facto AI referee: its scores cancelled a Hachette novel, embarrassed The New York Times and cast doubt on a major literary prize, and Substack now bakes the detector into its platform. Independent NBER benchmarking really does put it far ahead of GPTZero, Originality.ai and Turnitin on false positives — yet humanizer tools evade it over 96 percent of the time, verdicts wobble on short excerpts, and bias questions hang over who gets accused. This article weighs the evidence from WIRED’s investigation, the Chicago Booth benchmark and the company’s own disclosures to answer the practical question: when should you trust a Pangram score, and when should you not?

Read more
detect ai writing tells and limits a typewriter blank keys

Can You Teach Yourself to Detect AI Writing? Maybe — Here Is What the Evidence Actually Says

Can you teach yourself to detect AI writing? The honest answer is maybe. There is no single giveaway — only accumulating habits: signature vocabulary such as delve, meticulous and quietly, the “not X, but Y” frame, the rule of three, claim escalation and length that ignores the situation. The evidence is more encouraging than the old consensus: annotators who use these systems daily reached 86.7% to 96.7% true-positive rates individually and 99.3% as a majority vote across 300 articles, beating every automated detector except Pangram, while occasional users managed 56.7% — barely above chance. This breakdown covers the tells that hold up, the human-versus-detector numbers, the Stanford finding that seven detectors falsely flagged 61% of essays by non-native English writers, why the tells keep expiring, and what to do instead if you commission content.

Read more
CHAT