AI cheating is now common enough that some US educators have a name for it: “homework scams”. Yet the universities dealing with it are moving away from the software that promised to catch it. An AFP report published on 6 October, syndicated by the Bangkok Post, describes campuses from Yale to Cornell that have banned or discouraged relying on AI detection tools as the main evidence in misconduct cases.

The reason is simple to state and hard to fix. Detectors produce false positives, they have been shown to be biased against non-native English speakers, and a new crop of “humanising” tools can disguise AI text. Professors told AFP they are redesigning assessments instead, including pairing written work with oral exams.

This article looks at why universities are shunning AI detectors, the experiment that shows both the scale of AI cheating and the limits of tricks to catch it, the false-positive arithmetic, what faculty surveys say, and what employers and training providers in the UK can learn.

Why US Universities Are Shunning AI Detectors Over AI Cheating

ai cheating us universities shun ai detectors b lobster pot with a funnel entrance

AI chatbots can produce a credible answer to almost any written assignment within seconds. That has made large-scale AI cheating easier than ever. The obvious response was software that claims to tell human writing from machine writing. Many universities tried it, and many are now backing away.

What the universities say

The University of Wisconsin–Madison’s faculty guidelines, quoted by AFP, describe detection tools as “imperfect at best”, carrying “the risk of false positives”, “biased against non-native English speakers”, and unable to “prevent students from using these tools”. Its teaching centre’s academic integrity guidance says the university’s principles “discourage using so-called AI detection tools due to their unreliability and risk of false accusations”.

Cornell’s guidance for instructors is blunter: “Unfortunately, it is unlikely that detection technologies will provide a workable solution to this problem.” It adds that detection tools “cannot provide evidence” for their claims and that tests “have revealed significant margins of error”.

What a false accusation costs

A wrongful accusation of AI cheating is not a minor administrative error. AFP notes that false charges “can tarnish academic records or upend career prospects”, and that some students have sued universities that accused them. For an institution, a tool that is right most of the time can still be wrong about hundreds of students a year.

Why this is happening now

The detectors have not suddenly become worse. What has changed is the volume. As more students use chatbots, more work gets checked, and even a small error rate produces a steady flow of disputed cases. The arms race described later in this article makes the problem harder each term.

The Orange Trap: Measuring AI Cheating in One Class

ai cheating us universities shun ai detectors c wax tablet diptych with a stylus

The most vivid example in the AFP report comes from Timothy Paustian, a professor in the bacteriology department at UW–Madison. His experiment shows both how common AI cheating can be and why clever tricks do not last.

How the trap worked

Paustian set an assignment testing students’ ability to evaluate scientific arguments. He suspected some students would simply paste the instructions into a chatbot. One student had previously left in the phrase “I would be happy to help you with this research!”, a giveaway of chatbot text.

So he inserted a hidden instruction into the assignment telling any AI that read it to mention the colour orange in the middle of the third paragraph. Students reading the assignment normally would not see it. A chatbot given the whole text would follow it.

What it caught

Sixty of 350 students were caught. That is 17.1% of the class (60 divided by 350), or roughly one student in six, in a single assignment. The chart below shows the split.

Paustian’s hidden-prompt trap: students in a 350-person class (whole class = 100%)
Whole class 350 (100%)
Not caught by the trap 290 (82.9%)
Caught mentioning the colour orange 60 (17.1%)

The 17.1% is a floor, not an estimate of AI cheating in the class. Students who used AI but read the instructions first, or who used a chatbot that ignored the hidden line, would not appear in it.

Why the trick stopped working

AFP reports that today’s large language model chatbots ignore such prompts, and that more careful students now remove them. The trap worked once and is now, in the report’s word, redundant. Paustian’s own conclusion is stark: “I have pretty much thrown in the towel on classic writing assignments for lower-level classes.”

The cost to learning

He also named what is lost. “Students need to be taught how to think,” he told AFP. “Writing is a great way to demonstrate that, and LLMs are making that harder.” That is the real concern behind AI cheating. It is not only about grades, but about whether a written assignment still shows what a student can do.

What University Policies Now Say

ai cheating us universities shun ai detectors d bunsen burner under a lab tripod

The move away from detectors is not one central decision. Each university has reached it separately, which makes the pattern more striking.

InstitutionPosition on AI detectorsSource
UW–MadisonPrinciples discourage detection tools; suspected cases start with a meeting with the studentTeaching centre guidance
Cornell“Unlikely that detection technologies will provide a workable solution”Center for Teaching Innovation
YaleNamed by AFP among universities banning or discouraging reliance on detectorsAFP, 6 October 2026
VanderbiltDisabled Turnitin’s AI detector in August 2023Vanderbilt Brightspace guidance
Harvard CollegeDean proposes “getting out of the AI-detection business”Email to students, 2 September 2026

Vanderbilt went first

Vanderbilt’s August 2023 decision was one of the earliest. It said Turnitin had switched the feature on with less than 24 hours’ notice, no option to disable it at first and “no insight into how it works”. Three years later, its reasoning reads as the template many others have followed.

Harvard’s dean wants out

At Harvard, College Dean David Deming went further in an email to students at the start of term. According to The Harvard Crimson, he wrote: “We would all benefit from getting out of the AI-detection business because it damages trust between faculty and students.” He also suggested “AI acceptance, or even encouragement” for papers too large for a proctored exam.

Meeting the student first

UW–Madison’s guidance shows what replaces the software. If an instructor suspects AI cheating, they should “begin as you would with any other case of potential academic misconduct: by meeting with the student”. The conversation, not a score, becomes the starting point.

The False-Positive Arithmetic Behind AI Cheating Cases

ai cheating us universities shun ai detectors e overhead projector with a folding arm

The strongest argument against detectors is arithmetic. A tool that sounds accurate can still produce a large number of wrong accusations once it checks enough work.

Vanderbilt’s 750 papers

When Turnitin launched its detector, it claimed a false positive rate of 1%. Vanderbilt did the sum: it had submitted 75,000 papers to Turnitin in 2022, so “around 750 student papers could have been incorrectly labeled as having some of it written by AI”. That is 75,000 multiplied by 0.01, and every one of those 750 would be a student facing an AI cheating allegation they did not deserve.

The Stanford study on non-native writers

A 2023 Stanford study, “GPT detectors are biased against non-native English writers”, tested seven widely used detectors. They were close to perfect on essays by US eighth-graders. On 91 TOEFL essays written by non-native speakers, they misclassified more than half as AI-generated, with an average false positive rate of 61.22%.

91 human-written TOEFL essays checked by seven AI detectors (Liang et al., 2023)
Flagged as AI by at least one detector 89 of 91 (97.80%)
Average false positive rate across the seven detectors 61.22%
Flagged as AI by all seven detectors 18 of 91 (19.78%)

Every one of those essays was written by a person. The researchers found that the flagged essays had lower “perplexity”, a natural language processing measure of how predictable a text is. Simpler, more predictable word choices are common in writing by people still learning English.

The same text, made fancier, passes

The study’s most uncomfortable finding came next. When the researchers used a chatbot to “enhance the word choices to sound more like that of a native speaker”, the average false positive rate fell from 61.22% to 11.77%. In other words, polishing human writing with AI made it look more human to the detectors. That undermines the logic of using them as evidence of AI cheating.

We have looked at detector accuracy in more depth in our earlier piece, Spotting AI writing: how reliable are the detectors?

What Faculty Say About AI Cheating

ai cheating us universities shun ai detectors f coin spinning on its edge

Faculty attitudes help explain the shift. A survey by the American Association of Colleges and Universities and Elon University’s Imagining the Digital Future Center, carried out between 29 October and 26 November 2025, asked 1,057 US faculty about generative AI. The authors stress it is not a scientific sample and the results are not generalisable, but the numbers are striking.

Cheating has risen, faculty say

According to the full report, 78% said cheating on their campus has increased since generative AI tools became widely available, including 57% who said it has increased a lot. 73% said they had personally dealt with academic integrity issues involving students’ AI use. And 95% said AI would increase students’ overreliance on these tools, the figure AFP quoted.

Few faculty trust the detectors

The same survey shows why the detector market has stalled on campus. Only 31% of faculty said they use AI detection tools, even though 34% said their university provides a subscription. Just 23% rated the tools very or somewhat effective, while 33% said they were not very effective and 22% said not effective at all.

US faculty views on AI cheating and detection (AAC&U and Elon survey, 1,057 faculty)
Say cheating on campus has increased since generative AI 78%
Have personally dealt with an AI-related integrity issue 73%
Rate detection tools not very or not at all effective (33% + 22%) 55%
Actually use AI detection tools 31%

More than twice as many faculty think the tools fail (55%) as think they work (23%). That gap, more than any single policy, explains why AI cheating is being handled through assessment rather than software.

Confidence in their own judgement

The report adds a twist: a majority of faculty think their own ability to spot AI writing is adequate, but only 4% think their colleagues are very effective at it. That mismatch is another reason for caution. Human judgement about AI cheating is no more reliable than the software unless it is backed by a conversation and evidence.

Humanisers and the AI Cheating Arms Race

Even a perfect detector would face a moving target. A new category of “humanising” tools rewrites AI-generated text to make it read as human.

Detectors versus evaders

Misinformation researcher Timothy Caulfield described it to AFP as “a bit of a technology arms race happening: the detectors versus the evaders”. He added that “the ‘humanising’ programs, which make the writing seem even more authentic, are getting better and better”.

Why the arms race favours AI cheating

The economics are lopsided. A student needs to beat a detector once to submit an essay. A detector has to be right about every student, every time, without falsely accusing anyone. Each improvement in detection can be met by a rewrite, while each false accusation damages trust. That asymmetry is why institutions are looking for approaches that do not depend on winning the race.

What Replaces Detectors: Assessment Redesign

If AI cheating cannot be reliably detected after the fact, the answer is to design assessments where it is harder to do or easier to see.

Oral exams and conversations

Four professors told AFP they are rethinking how they evaluate students, including pairing take-home written work with oral exams. A short conversation about a submitted essay quickly shows whether a student understands what they handed in. Paustian is considering “some sort of video debate”, though he warns that students could still read AI-generated content aloud.

Proctored exams for mastery

Deming’s email at Harvard argued that “take-home exams and problem sets are no longer a sustainable way to assess mastery”, and proposed shifting towards proctored exams. Harvard’s Bok Center had already advised faculty to give more weight to in-person assessment.

AI logs and open use

Some instructors are taking the opposite approach and allowing AI openly but asking students to document it. The Crimson reported that one Harvard professor has students keep “AI logs” recording how they used AI to prepare. When use is declared, the question shifts from catching AI cheating to judging the quality of the student’s own thinking.

Process over product

A broader pattern is to grade the process as well as the result: drafts, notes, sources and in-class writing. This makes it harder to submit work produced entirely by a chatbot and gives instructors evidence that does not rely on a detector score.

The Harvard Split Over AI in Writing

Deming’s proposal has divided his own faculty, which shows that there is no easy consensus on how to respond to AI cheating.

What the courses allow

A Crimson analysis of a random sample of roughly 600 fall 2026 courses found that more than half of science and engineering courses allowed some use of AI. That fell to 35% in the Social Sciences Division and 27% in the Arts and Humanities Division.

Faculty push back

English professor Deirdre Lynch told the Crimson she did not want to be “the cop” checking assignments for AI, but would keep a blanket ban in her courses. Historian Peter Gordon warned that AI in writing classes could push students to outsource their thinking. Most writing faculty agreed they wanted to stop policing AI use, but were divided on encouraging it.

What it tells us

The split is useful for anyone outside academia. Avoiding detection software is becoming the consensus, but what replaces it depends on what an assessment is meant to prove. A coding course and a poetry seminar need different rules.

Lessons for UK Employers and Training Providers

AI cheating is not only a university problem. Any organisation that judges people by written work faces the same issue, from recruitment tasks to professional exams.

Recruitment take-home tasks

Many employers set written or coding take-home tasks for candidates. These have the same weakness as take-home exams. Following the universities, pair them with a short live discussion of the submitted work, which reveals understanding far better than a detector score.

Do not use detector scores as evidence

If your HR or compliance team uses AI detection, treat the score as a prompt for a conversation, never as proof. The Vanderbilt arithmetic applies to any organisation that checks large volumes of writing, and the Stanford findings raise discrimination risks for staff and candidates who are not native English speakers.

Set clear AI use rules

The universities moving away from detectors are replacing them with clear expectations about when AI may be used. Businesses need the same: a written policy on acceptable AI use for tasks, training and assessments. Our AI strategy team helps organisations write these, and our IT governance work covers how to enforce them.

Design for understanding

The lasting lesson from the campus debate is that the best defence against AI cheating is an assessment that tests understanding directly. Ask people to explain, apply and defend their work, and the question of who wrote the first draft matters much less.

AI Cheating and AI Detectors: Frequently Asked Questions

Why are universities stopping using AI detectors?

Because they produce false positives, have been shown to be biased against non-native English writers, and can be defeated by “humanising” tools. Universities including UW–Madison and Cornell say they are not reliable enough to use as evidence of AI cheating.

How common is AI cheating?

There is no reliable national figure. In one UW–Madison class, a hidden prompt caught 60 of 350 students, or 17.1%, which is a minimum for that assignment. In a 2025 survey, 78% of US faculty said cheating had increased since generative AI became widely available.

How accurate are AI detectors?

It varies by tool and by writer. A 2023 Stanford study found seven detectors wrongly flagged an average of 61.22% of essays by non-native English writers as AI-generated, while being close to perfect on essays by US students.

What are universities doing instead?

Redesigning assessment: oral exams, proctored exams, in-class writing, grading drafts and process, and in some courses allowing AI openly with a log of how it was used.

Should businesses use AI detectors on job applications?

Not as evidence. A detector score should at most prompt a conversation with the candidate, and live discussion of their work is a more reliable test.

References