AI vision systems are very good at saying where an object is in a picture. A new study from York University in Toronto shows that they are not good at something people do without thinking: letting what they have just seen change what they see next. Using a classic optical illusion, the researchers found that human observers and neurons in a monkey’s visual cortex shift their sense of where an object is, while every AI vision model they tested stayed fixed on the raw pixels.

That might sound like a point in the machines’ favour. The pixels did not move, so why should the answer change? But the researchers argue that this “mistake” is a feature of biological vision, part of how brains make sense of a moving world. If AI vision is meant to work alongside people, and to see the world in ways people can predict, it may need some of the same quirks.

This article explains the illusion, how the experiment worked, what the brain and the models did, why it matters for anyone building or buying machine vision, and what the study does not show.

What the York Study Found About AI Vision

ai vision visual illusion motion aftereffect b barbers pole with spiral stripes turning

The paper, titled “The macaque IT cortex but not current artificial vision networks encode object position in perceptually aligned coordinates”, is published in Current Biology. A preprint has been on arXiv since 11 March 2026.

The question

The researchers asked whether object position signals in the brain reflect what we perceive or simply where the object sits on the retina. In ordinary images the two are the same, so you cannot tell them apart. An illusion that moves perceived position without moving the image solves that problem. The second question was whether today’s AI vision networks behave like the brain under the same illusion.

The answer in one line

People and monkey brain cells shifted their estimate of where an object was after watching motion. Feedforward, recurrent and video-based AI vision networks did not, even though they encoded object position accurately.

Who did the work

The first author is Elizaveta Yakubovskaya, a York graduate student. The co-authors are Hamidreza Ramezanpour, Matteo Dunnhofer, who is also affiliated with the University of Udine in Italy, and senior author Kohitij Kar, an assistant professor at York and Canada Research Chair in Visual Neuroscience. Kar is a member of York’s Centre for Vision Research and the Connected Minds programme.

The Motion Aftereffect, Explained

ai vision visual illusion motion aftereffect c model eyeball with a camera shutter iris

The illusion at the centre of the study is the motion aftereffect. Its best-known form is the waterfall illusion, described in 1834 by Robert Addams after he stared at the Falls of Foyers in Scotland and then looked at the rocks beside it, which seemed to drift upwards.

What happens

If you stare at something moving steadily in one direction for a while, then look at something still, the still object briefly appears to move, or to sit slightly displaced, in the opposite direction. Nothing in the image has changed. Your visual system has adapted to the motion and is briefly out of balance.

Position, not just motion

Most people know the aftereffect as illusory movement. This study used a related effect: a shift in perceived position. After adapting to rightward motion, people judged a still object to be a little to the left of where it really was, and vice versa.

Why scientists use it

The illusion gives researchers a clean test. The pixels are identical before and after adaptation, so any system that reports a change must be using recent visual history, not just the current image. That makes it a sharp tool for comparing brains with AI vision models.

How the Experiment Worked

ai vision visual illusion motion aftereffect d macaque monkey sitting on a rock

The team combined three kinds of evidence: human perception, recordings from monkey brains and tests on artificial networks.

PartWho or whatMethod
Baseline localisation35 human participantsClick on the centre of an object seen for 100 ms
Illusion test22 participants per motion direction30 s of drifting stripes, 3 s top-ups, then the same task
Brain recordingsTwo rhesus macaquesTwo or three 96-electrode arrays in IT cortex; 3 s of motion, then a still image
Model testsImage and video networksSame images and motion; position read out from internal features

The images

The images showed eight kinds of object, a bear, an elephant, a face, a car, a dog, an apple, a chair and an aeroplane, placed at different positions against natural-looking backgrounds. People were very consistent about where each object was, which gave the team a reliable benchmark.

The brain area

The recordings came from the inferior temporal (IT) cortex. IT sits late in the brain’s “what” pathway and is best known for recognising objects. Earlier work had shown that object position can also be read from IT activity. The open question was whether that position signal matches perception or just mirrors the image.

Reading position from activity

For both brains and machines, the team trained simple linear decoders to read an object’s horizontal and vertical position from activity patterns. The decoders were trained on responses to still images with no prior motion, then applied, without retraining, to responses recorded after motion adaptation.

What People and the Brain Did

ai vision visual illusion motion aftereffect e spiral disc spinning on a stand

The human results confirmed the illusion, and the monkey recordings showed a matching shift in the brain.

Humans shifted opposite the motion

After rightward motion, people placed objects on average 0.20 degrees of visual angle further left. After leftward motion, they placed them 0.13 degrees further right. Vertical estimates did not change, as expected for a horizontal illusion, and the estimates stayed just as consistent. Adaptation biased perception without making it noisier.

IT cortex shifted the same way

Position read from IT activity moved in the same direction as human perception. After rightward motion, decoded positions shifted left by an average of 0.71 degrees, a statistically significant change. After leftward motion, the shift to the right was smaller and not statistically significant on its own, but the difference between the two conditions had the same direction as in people.

The brain’s internal map changed shape

Motion adaptation did more than turn activity down. It reshaped the pattern of activity across the neural population. Using a similarity measure called centred kernel alignment, the team found that the population’s representation after adaptation was markedly less similar to its baseline than noise alone would explain.

What the AI Vision Models Did

ai vision visual illusion motion aftereffect f garden gnome beside a shifted ghost of itself

This is where the gap appeared. The researchers tested several well-known AI vision architectures, from classic image classifiers to models trained on video.

SystemTypeIllusion shift, degrees
Human observersPerception0.33
Macaque IT cortexNeural recordings0.84
SlowFastVideo action recognition0.019
ResNet-18Image classifier0.0018
SimCLR ResNet-50Self-supervised image model0.0012
VGG-16Image classifier0.00058
ViT-L/32Vision transformer0.00032
AlexNetImage classifier-0.00022
ConvRNNRecurrent video model-0.0016

The image models in this table were tested with simulated adaptation, activity decay fitted to real IT neurons; the video models received drifting-stripe movies. Values near zero, positive or negative, mean no meaningful shift.

They know where things are

The models were not short of position information. When the team read position from their features in ordinary conditions, VGG-16 scored a correlation of 0.84 with true horizontal position and ResNet-18 scored 0.81, against about 0.60 for IT activity with more than 150 recorded units. The transformer ViT-L/32 was weaker at 0.59. AI vision networks can match or beat the brain on this basic task.

How accurately position can be read out in normal viewing (correlation with true horizontal position, from the paper)

VGG-16: 0.84
ResNet-18: 0.81
Macaque IT, over 150 units: about 0.60
ViT-L/32: 0.59

But they do not adapt

Standard feedforward networks have no memory of the previous frame, so by design they give the same answer before and after adaptation. The team therefore built in “fatigue”: unit activity that decays over time, using time constants measured from the real IT neurons. Activity dropped as expected, but the position estimates still did not shift in the illusion’s direction.

Video models did not adapt either

The obvious next candidates were models built for motion. SlowFast, an action recognition network, and ConvRNN, which has recurrent connections and feedback, were shown adapter movies of drifting stripes followed by still images. Neither produced a meaningful shift. Temporal processing, as currently built into these AI vision models, was not enough.

Size of the illusion shift in degrees (bar length relative to macaque IT, 0.84 = 100%)

Macaque IT cortex: 0.84
Human observers: 0.33
SlowFast video model: 0.019
ResNet-18: 0.0018
VGG-16: 0.00058

The neuralization test

The team then tried a shortcut. They measured how motion adaptation transformed IT activity and applied that same transformation to the features of VGG-16, ResNet-18 and SimCLR, a process they call “neuralization”. Once transformed, the models showed the same direction-opposite shifts as people and monkeys. The capacity is there; what is missing is a mechanism that produces the transformation by itself.

Why a Perceptual "Mistake" Can Be a Feature

It is tempting to read the result as AI vision beating human vision: the machine reports the true position, the person is fooled. The researchers see it differently.

Brains are not cameras

Biological vision does not just log pixel coordinates. It constantly adjusts to recent experience, which helps it stay sensitive to change and efficient with energy. The motion aftereffect is a side effect of that adjustment. Neuroscience News, summarising the paper, described the brain’s shifts as “an adaptive feature rather than a flaw”.

Getting the right answer is not enough

“If we want AI that works with humans and understands the world in more human-compatible ways, we cannot focus only on whether it gets the right answer,” Kar said in York’s announcement. “We also need to understand the computations that produce human perception and behavior.”

The NeuroAI argument

Kar described the study as capturing “the promise of NeuroAI”, using careful experiments to find computations that brains use and AI still lacks, then building them into better systems. On this view, illusions are not party tricks. They are probes that show which parts of perception a model has, and which it does not.

Why the AI Vision Gap Matters Outside the Lab

The experiment is narrow, but the question behind it is not. Whenever people and machines look at the same scene and need to agree, differences in how they see can matter.

Driving and robotics

A driver who has just watched traffic flow past, and a camera system that has not adapted at all, may judge a stationary object’s position slightly differently. In most cases the gap is tiny. In shared-control situations, where a person takes over from an automated system, understanding how each “sees” is part of designing a safe handover.

Medical imaging and inspection

Clinicians and inspectors scroll through moving or sequential images. If an AI vision assistant and a human reviewer are biased in different ways by recent images, they may disagree in ways that are hard to explain. Knowing that the model has no history effects, while the human does, is useful context for anyone designing those workflows.

Better benchmarks

The study’s authors say their results provide a new benchmark for models of dynamic vision. Most benchmarks reward models for matching ground truth. This one asks whether a model matches perception, which is a different and arguably harder test for computer vision.

Where motion-aware hardware fits

Other researchers are tackling time and motion from the hardware side. A team at KAIST recently built a stacked AI chip with fast and slow layers to recognise changes in motion over time. Approaches like that build history into the system, which is the ingredient the York study found missing.

What Teams Using AI Vision Should Take From This

Most organisations will never train a vision model from scratch. They buy one inside a camera system, an inspection tool or a document-processing service. The study still has practical lessons for them.

Test with sequences, not just single frames

Many evaluations score a model on still images. Real deployments see video and streams of images, where context builds up. If your use case involves motion, such as people moving through a site or products on a conveyor, test AI vision on realistic sequences and check whether recent frames change its answers.

Expect humans and machines to disagree differently

A human reviewer’s judgement is shaped by what they have just looked at. An AI vision model’s judgement, on this evidence, is not. Where both review the same material, build a process for resolving disagreements rather than assuming one side is always right.

Ask vendors about temporal behaviour

When buying an AI vision product, ask how it uses previous frames, whether it smooths or tracks objects over time, and how it was tested on video. Vendors that can answer clearly have usually thought about the problem.

Do not over-read one study

The York result is about a specific illusion and a specific set of research models. It does not mean commercial AI vision systems are unsafe. It means there is a measurable difference between how they and people process motion, and that difference is worth knowing about.

Try the Illusion Yourself

You do not need a laboratory to experience the effect the York team used to test AI vision. A minute and a moving scene are enough.

Find steady motion

A waterfall, a river, a scrolling credits sequence or a video of falling snow all work. The motion should be steady and fill a good part of your view.

Stare for about 30 seconds

Fix your eyes on one point in the middle of the motion and keep them still. The human participants in the study adapted for 30 seconds before their first test.

Look at something still

Switch your gaze to a still surface, such as a wall or a patterned rug. For a few seconds it will seem to drift the other way, or objects will seem slightly shifted. That is your visual system recalibrating, and it is exactly what the AI vision networks in the study did not do.

What the Study Does Not Show

The result is clear, but its scope is limited, and the paper is careful about that.

A small number of animals and one illusion

The brain recordings come from two monkeys, and the study uses one illusion in one direction of motion. The IT shift after leftward adaptation was not significant on its own. Broader tests, with other illusions and more subjects, would strengthen the conclusion.

Not the latest multimodal systems

The models tested are well-studied research architectures, including AlexNet, VGG-16, ResNet-18, ViT-L/32, SimCLR, SlowFast and ConvRNN. The paper does not test the large multimodal assistants that now answer questions about images. It would be wrong to assume they behave differently, but it would also be wrong to claim the study tested them.

IT is not the final step

The authors note that the shift in IT came with a drop in consistency that people did not show, which suggests areas downstream of IT add further processing before behaviour. IT is a likely place where perceptually aligned position is represented, not necessarily where the decision is made.

The fix was imposed from outside

Neuralization shows that existing AI vision features can express the illusion, but the transformation came from monkey data. Nobody has yet built a model that learns to produce it on its own.

What Comes Next for AI Vision Research

The paper ends with a research agenda for both neuroscience and AI vision engineering.

Mapping the brain circuit

The authors suggest recording from IT together with motion areas such as MT and V4, temporarily silencing those areas during adaptation, and recording across cortical layers to separate feedforward from feedback signals. That would show where the history effect comes from.

New model designs

The paper argues that the missing ingredient may be interaction between motion-selective and form-selective channels, combined with history-dependent gain control. That points towards architectures that explicitly couple form and motion pathways.

New training goals

Today’s AI vision models are mostly trained to label images or actions correctly. The authors suggest objectives that reward perceptual alignment, not just correct labels. That would be a significant change in how vision models are built and evaluated.

From lab to products

Turning a neuroscience finding into a better product is slow, as our piece on why translating AI advances to the real world is hard explains. Expect this to show up first in research benchmarks, then in specialised systems, long before it reaches consumer apps.

AI Vision and the Illusion: FAQ

What did the York study find?

People and monkey brain cells shifted their sense of an object’s position after watching motion, as the motion aftereffect predicts. The AI vision networks tested did not shift, even though they encoded position accurately.

What is the motion aftereffect?

It is an illusion in which, after watching steady motion, a still object seems to move or sit displaced in the opposite direction. The waterfall illusion is the classic example.

Which AI vision models were tested?

AlexNet, VGG-16, ResNet-18, ViT-L/32, SimCLR ResNet-50, and the video models SlowFast and ConvRNN.

Does this mean AI vision is worse than human vision?

Not exactly. The models were accurate about position. The difference is that they do not adapt to recent experience, which people and primates do. That matters most when humans and machines need to see the world the same way.

Where was the study published?

In Current Biology, by Elizaveta Yakubovskaya, Hamidreza Ramezanpour, Matteo Dunnhofer and Kohitij Kar. A preprint is on arXiv.

Could AI vision models learn this?

The study shows that their features can express the effect if an IT-like transformation is applied. Building a model that produces it on its own is the next challenge.

References and Further Reading