AI summaries that get a single detail wrong can change what people remember about an event they watched with their own eyes. That is the central finding of a new study from Georgetown University and the University of Washington, which showed participants a short video of a car hitting a pedestrian and then gave them a ChatGPT-written summary of it.

When the summary described the traffic sign correctly, 83.6% of people later remembered the sign correctly. When the summary named the wrong sign, only 44.8% did. Telling people the summary came from AI made no measurable difference, and neither did how much they trusted AI.

The same paper found that the AI summaries themselves were riddled with errors. Every one of the 20 AI summaries the researchers generated with ChatGPT and Gemini left out important facts, and 95% missed the collision itself. This article explains how the study worked, what each number means, why the result fits five decades of memory research, and what organisations using AI to summarise video, calls or documents should change.

What the AI Summaries Study Found

ai summaries misleading distort human memory b octagonal stop sign on a post

The paper, “AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory”, was posted to arXiv on 23 September 2026. A shorter version is published in the proceedings of the Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES), taking place in October 2026.

The headline result

Two findings carry the paper. First, AI summaries of simple videos contained frequent errors, above all omissions of central details. Second, a misleading AI summary distorted people’s memory of the original event, and it did so whether they believed a machine or a person had written it. In the authors’ words, “human memory can instead be distorted by these mistakes.”

Who ran it

The study was led by Mattea Sim, an assistant research professor at the Massive Data Institute in Georgetown’s McCourt School of Public Policy. Her co-authors were Yael Eiger, a PhD candidate at the University of Washington, and Tadayoshi Kohno, the McDevitt Chair in Computer Science, Ethics, and Society at Georgetown. The data and materials are public on the Open Science Framework.

“AI is a new method of delivering misinformation, and it has the potential to create these false memories for people who are reading that information,” Sim said in Georgetown’s announcement.

How the Researchers Tested AI Summaries

ai summaries misleading distort human memory c stack of paper sheets with a folded corner

The design had two parts. One measured how often AI summaries go wrong. The other measured what a wrong summary does to a reader’s memory.

Part one: grading the machines

The team asked two widely used models with video-to-text features, ChatGPT-5.5 and Gemini 2.5 Flash-Lite, to produce AI summaries of two 25-second animated videos. Video summarisation of this kind combines computer vision with language generation, so it has two places to go wrong. Each model summarised each video five times, from a fresh upload and a fresh session, giving 20 summaries in all.

The prompt was deliberately generous. It asked for “a factual, neutral text summary that is at least 300 words”, told the model not to “introduce new facts, speculation, or interpretations”, and to “make sure everything important is included.” Three researchers agreed a list of 19 central details, and two coded all the AI summaries against it.

Part two: a classic memory experiment, updated

The videos come from a 2026 replication of a famous 1978 experiment by Elizabeth Loftus and colleagues. A red car turns right at a junction marked either by a stop sign or by a yield sign, hits a pedestrian, and the two people then meet in the road. Participants watched one version, then came back 24 to 48 hours later to read a 21-sentence ChatGPT summary before a 12-question memory test.

The summary was identical for everyone except one detail. Half read the correct sign. The other half read the wrong one. Separately, half were told an AI transcription tool wrote the summary and half were told a professional human transcriber did.

Who took part

The researchers recruited 360 US adults on Prolific. Five failed an early attention check, 331 finished both parts, and 328 were analysed after three more failed a second check. The four groups ranged from 80 to 83 people each, above the 63 per group the power analysis required.

GroupTold the summary was written byTraffic sign in the summaryParticipants
AI, consistentAn AI transcription toolCorrect82
AI, misleadingAn AI transcription toolWrong80
Human, consistentA professional transcriberCorrect83
Human, misleadingA professional transcriberWrong83

AI Summaries Got Most of the Story Wrong

ai summaries misleading distort human memory d small sedan car

The first half of the paper is a short audit, not a large benchmark, but its results are stark. All 20 AI summaries contained mistakes, with at least 7 errors and as many as 21 in texts that averaged just 305 words.

Share of summaries with at least oneChatGPTGeminiAll 20
Omission of a central detail100%100%100%
Critical inaccuracy30%100%65%
Non-critical inaccuracy80%100%90%
Addition (hallucinated detail)70%50%60%

Omission was universal

Every one of the AI summaries left something important out. Across both models, the AI summaries omitted 51.6% of the 19 central details on average, roughly ten per summary. All but one missed the single most important event, the car striking the pedestrian. One summary went as far as saying “no other characters or vehicles interacted directly with the scene.”

Central details (of 19)OmittedDescribed wronglyDescribed correctly
ChatGPT52.1%3.2%44.7%
Gemini51.1%9.0%40.0%
Both models51.6%6.1%42.4%

Wrong details and invented ones

Beyond omissions, the AI summaries averaged 1.15 errors in central details, 2.05 errors in peripheral details such as sign text or car positions, and 1.6 hallucinated additions. Summaries invented extra vehicles and pedestrians, erratic driving, and guesses about what people intended. One stated that the pedestrian “does not enter the roadway”, which reverses the event that caused the accident.

Share of the 20 AI summaries containing each type of error
Omission of a central detail 100%
Non-critical inaccuracy 90%
Critical inaccuracy 65%
Hallucinated addition 60%

Errors changed from run to run

The mistakes in the AI summaries were not consistent. A car appeared in one summary and vanished from the next, and the main car turned right in one and left in another. That matters in practice: an organisation cannot test a summariser once, note its quirks and correct for them, because the next run may fail differently. ChatGPT added more invented details, while Gemini made more errors describing central events.

“I was struck by how bad the summaries were, even at this stage in AI development,” Eiger said.

A Misleading Summary Rewrote What People Remembered

ai summaries misleading distort human memory e video camera on a tripod

The memory test turned on one question: did the car stop at a stop sign or a yield sign? People had seen the answer in the video a day or two earlier. The only thing that differed was what the summary said.

Participants who correctly remembered the traffic sign, by summary condition
Read an accurate summary 83.6%
Watched the video only, earlier replication (different sample) 80.3%
Misleading summary labelled as AI 46.3%
Read a misleading summary, all labels 44.8%
Misleading summary labelled as human 43.4%

83.6% versus 44.8%

The gap is 38.8 percentage points, and it was statistically strong. Put the other way round, wrong answers rose from 16.4% to 55.2%, so a single false detail in the summary more than tripled the error rate. Because two options were offered, 44.8% is below the 50% a coin toss would score: most people who read the misleading version now remembered the sign the summary named, not the one they saw.

Labelling the summary as AI did not help

Among people who read the misleading version, 46.3% of those told it was AI-written answered correctly, against 43.4% of those told a human wrote it. The difference was not significant. A clear AI label, the kind of disclosure many policies rely on, did not protect memory.

Trust in AI made no difference

The researchers also measured trust in AI on a 10-item scale and asked how often people used chatbots. Neither changed how susceptible they were. Sceptics were misled as often as enthusiasts. “An awareness that AI makes mistakes may not be enough,” the authors write, and they suggest that AI literacy on its own may not counter the effect.

People still remembered the video

The result is not simply people answering from the text. On a question the summary never mentioned, whether the accident happened by day or at night, 96.6% answered correctly, and participants averaged 91.4% across the other non-critical questions. They remembered the video well. The misleading line changed one specific memory.

Why AI Summaries Fit a 50-Year-Old Memory Problem

ai summaries misleading distort human memory f star shaped badge on a stand

Psychologists have studied this pattern since the 1970s under the name the misinformation effect. Information received after an event can alter memory of the event itself.

The misinformation effect

In the original 1978 stop-sign study, people exposed to misleading information were far more likely to misremember the sign. Later work showed that a single word matters: asking how fast cars “smashed” rather than “hit” changed speed estimates. In “rich false memory” studies, 25% to 35% of adults came to recall being lost in a shopping mall as a child. False memories often persist even after people are warned about the misinformation. The new paper shows that AI summaries can act as the misleading source, even with no bad intent behind them.

Omissions are not neutral

The most common AI error was leaving things out, and the authors argue that this is not harmless either. In classic studies, people given no information about a detail recalled it less accurately than people given the correct information. A summary that skips the collision may not plant a false memory, but it can weaken the true one. Hallucinated additions resemble “false presuppositions”, which earlier research found can implant memories of objects that were never there.

Where AI Summaries Are Already Used in High-Stakes Work

The authors chose a traffic accident on purpose. Police departments are among the fastest adopters of AI report writing, and eyewitness memory has always been the misinformation effect’s home ground.

SettingWhat gets summarisedWho might be misled
PolicingBody-camera footage and audio turned into draft reportsOfficers, prosecutors, witnesses
HealthcareConsultations turned into clinical notesClinicians and patients
WorkplacesMeeting transcripts turned into minutes and actionsAttendees recalling decisions
News and searchArticles turned into previews and overviewsReaders who skip the source
Legal workDepositions, evidence and case filesLawyers, judges, juries

Police reports

Axon’s Draft One generates police report drafts from body-camera audio, and similar tools are spreading. The paper cites a case in Heber City, Utah, where an AI-drafted report said an officer had turned into a frog, because the Disney film “The Princess and the Frog” was playing in the background. That error was easy to spot. The study’s warning is about subtler errors that no one notices and that can then reshape the officer’s own recollection.

Legislators have started to respond. Utah’s SB 180 requires a disclosure on AI-assisted police reports and an officer’s certification of accuracy. California’s SB 524 goes further, requiring a disclaimer, an audit trail and retention of the original AI draft. The study suggests disclosure alone will not stop the memory effect, which makes the draft-retention rule the more important of the two.

Healthcare, meetings and news previews

The same risk applies wherever people read AI summaries of something they also experienced, such as a consultation, a meeting or a call. Eiger noted the team had no access to enterprise versions of these tools and called for more audits. Separate research has shown chatbots drifting into misinformation over long AI conversations, and courts have already sanctioned lawyers over AI-hallucinated witnesses.

What Organisations Should Change About AI Summaries

The usual safeguard, a human in the loop, assumes the human’s memory is a fixed reference that can catch the machine’s mistakes. This study shows the reference can move. Controls therefore need to protect the original record and the order in which people meet it.

ControlWhy it helpsEvidence from the study
Record your own account firstCaptures memory before the summary can alter itMisleading text changed recall a day later
Check the summary against the sourceCatches errors instead of absorbing themAll 20 summaries had errors
Checklist of required detailsMakes omissions visible95% omitted the collision
Keep the AI draft and the sourceAllows later audit of what changedErrors varied from run to run
Do not rely on labels or training aloneNeither prevented the effectAI labels and trust levels made no difference

Record your own account first

Where a person witnessed the event, as in a call, a meeting or an incident, ask them to write their own account before they read the AI summary. Once the summary has been read, there is no clean way to separate what they remember from what they were told. This is the cheapest control and the one most directly supported by the study.

Design AI summaries to show gaps

AI summaries that look complete invite trust. Tools that produce AI summaries can list the fixed fields they were asked to cover and mark any they could not fill, so a missing collision shows up as an empty field rather than a silence. For recurring AI summaries, a short checklist of required details, agreed in advance, turns omissions into something a reviewer can see.

Keep the draft and the source

Retain the original recording and the unedited AI draft alongside the final version, as California now requires for police reports. That creates an audit trail if a dispute arises later about who remembered what. It also lets teams measure how often their summariser makes each kind of error, which is the only way to know whether it is fit for purpose. Our IT governance and AI strategy work treats this kind of record-keeping as a baseline for any AI tool that touches evidence.

Do not rely on warnings alone

Warnings given before people read AI summaries can help them notice discrepancies, but the authors point out that earlier research found warnings often fail. In this study, neither an AI label nor scepticism about AI protected memory. The researchers go further, suggesting the deeper question is “whether AI should be used in these contexts at all.”

Limits of the AI Summaries Study

The authors are careful about what their work does and does not show, and readers should be too.

Animated videos and one prompt

The audit used two short animated clips made in Grand Theft Auto V, one prompt and 20 summaries. That is a small sample of AI summaries, and real footage is messier. Against that, the authors argue the test was conservative: simple videos and a prompt demanding completeness gave the models their best chance, and other research on longer, more realistic videos has found similar errors.

One kind of misinformation and consumer models

The memory experiment planted one type of error in the AI summaries, a wrong central detail, and did not test what repeated omissions or several small errors do together. The summary shown to participants was itself chosen after repeated re-prompting, because ChatGPT often omitted the accident. The models tested were consumer versions, not the enterprise products that police forces or hospitals buy. None of that weakens the core finding. It marks out where the next studies need to go.

AI Summaries FAQs

Can AI summaries really change what I remember?

In this study, yes. People who read a summary with one wrong detail were far less likely to remember that detail correctly, 44.8% against 83.6% for an accurate summary.

Does knowing a summary is AI-generated protect me?

Not in this experiment. People told the summary was AI-written were misled about as often as those told a human wrote it.

Which AI models were tested?

ChatGPT-5.5 and Gemini 2.5 Flash-Lite for the error audit. The summary shown to participants was generated by ChatGPT.

What was the most common error in AI summaries?

Leaving things out. Every summary omitted central details, and 95% omitted the collision at the centre of the video.

What should teams do before using AI summaries?

Have people record their own account first, check summaries against the source, keep the original draft and recording, and use a checklist so omissions are visible.

References