AI vision systems are very good at saying where an object is in a picture. A new study from York University in Toronto shows that they are not good at something people do without thinking: letting what they have just seen change what they see next. Using a classic optical illusion, the researchers found that human observers and neurons in a monkey’s visual cortex shift their sense of where an object is, while every AI vision model they tested stayed fixed on the raw pixels.
That might sound like a point in the machines’ favour. The pixels did not move, so why should the answer change? But the researchers argue that this “mistake” is a feature of biological vision, part of how brains make sense of a moving world. If AI vision is meant to work alongside people, and to see the world in ways people can predict, it may need some of the same quirks.
This article explains the illusion, how the experiment worked, what the brain and the models did, why it matters for anyone building or buying machine vision, and what the study does not show.
Table of contents
- What the York Study Found About AI Vision
- The Motion Aftereffect, Explained
- How the Experiment Worked
- What People and the Brain Did
- What the AI Vision Models Did
- Why a Perceptual “Mistake” Can Be a Feature
- Why the AI Vision Gap Matters Outside the Lab
- What Teams Using AI Vision Should Take From This
- Try the Illusion Yourself
- What the Study Does Not Show
- What Comes Next for AI Vision Research
- AI Vision and the Illusion: FAQ
- References and Further Reading
What the York Study Found About AI Vision
The paper, titled “The macaque IT cortex but not current artificial vision networks encode object position in perceptually aligned coordinates”, is published in Current Biology. A preprint has been on arXiv since 11 March 2026.
The question
The researchers asked whether object position signals in the brain reflect what we perceive or simply where the object sits on the retina. In ordinary images the two are the same, so you cannot tell them apart. An illusion that moves perceived position without moving the image solves that problem. The second question was whether today’s AI vision networks behave like the brain under the same illusion.
The answer in one line
People and monkey brain cells shifted their estimate of where an object was after watching motion. Feedforward, recurrent and video-based AI vision networks did not, even though they encoded object position accurately.
Who did the work
The first author is Elizaveta Yakubovskaya, a York graduate student. The co-authors are Hamidreza Ramezanpour, Matteo Dunnhofer, who is also affiliated with the University of Udine in Italy, and senior author Kohitij Kar, an assistant professor at York and Canada Research Chair in Visual Neuroscience. Kar is a member of York’s Centre for Vision Research and the Connected Minds programme.
The Motion Aftereffect, Explained
The illusion at the centre of the study is the motion aftereffect. Its best-known form is the waterfall illusion, described in 1834 by Robert Addams after he stared at the Falls of Foyers in Scotland and then looked at the rocks beside it, which seemed to drift upwards.
What happens
If you stare at something moving steadily in one direction for a while, then look at something still, the still object briefly appears to move, or to sit slightly displaced, in the opposite direction. Nothing in the image has changed. Your visual system has adapted to the motion and is briefly out of balance.
Position, not just motion
Most people know the aftereffect as illusory movement. This study used a related effect: a shift in perceived position. After adapting to rightward motion, people judged a still object to be a little to the left of where it really was, and vice versa.
Why scientists use it
The illusion gives researchers a clean test. The pixels are identical before and after adaptation, so any system that reports a change must be using recent visual history, not just the current image. That makes it a sharp tool for comparing brains with AI vision models.
How the Experiment Worked
The team combined three kinds of evidence: human perception, recordings from monkey brains and tests on artificial networks.
| Part | Who or what | Method |
|---|---|---|
| Baseline localisation | 35 human participants | Click on the centre of an object seen for 100 ms |
| Illusion test | 22 participants per motion direction | 30 s of drifting stripes, 3 s top-ups, then the same task |
| Brain recordings | Two rhesus macaques | Two or three 96-electrode arrays in IT cortex; 3 s of motion, then a still image |
| Model tests | Image and video networks | Same images and motion; position read out from internal features |
The images
The images showed eight kinds of object, a bear, an elephant, a face, a car, a dog, an apple, a chair and an aeroplane, placed at different positions against natural-looking backgrounds. People were very consistent about where each object was, which gave the team a reliable benchmark.
The brain area
The recordings came from the inferior temporal (IT) cortex. IT sits late in the brain’s “what” pathway and is best known for recognising objects. Earlier work had shown that object position can also be read from IT activity. The open question was whether that position signal matches perception or just mirrors the image.
Reading position from activity
For both brains and machines, the team trained simple linear decoders to read an object’s horizontal and vertical position from activity patterns. The decoders were trained on responses to still images with no prior motion, then applied, without retraining, to responses recorded after motion adaptation.
What People and the Brain Did
The human results confirmed the illusion, and the monkey recordings showed a matching shift in the brain.
Humans shifted opposite the motion
After rightward motion, people placed objects on average 0.20 degrees of visual angle further left. After leftward motion, they placed them 0.13 degrees further right. Vertical estimates did not change, as expected for a horizontal illusion, and the estimates stayed just as consistent. Adaptation biased perception without making it noisier.
IT cortex shifted the same way
Position read from IT activity moved in the same direction as human perception. After rightward motion, decoded positions shifted left by an average of 0.71 degrees, a statistically significant change. After leftward motion, the shift to the right was smaller and not statistically significant on its own, but the difference between the two conditions had the same direction as in people.
The brain’s internal map changed shape
Motion adaptation did more than turn activity down. It reshaped the pattern of activity across the neural population. Using a similarity measure called centred kernel alignment, the team found that the population’s representation after adaptation was markedly less similar to its baseline than noise alone would explain.
What the AI Vision Models Did
This is where the gap appeared. The researchers tested several well-known AI vision architectures, from classic image classifiers to models trained on video.
| System | Type | Illusion shift, degrees |
|---|---|---|
| Human observers | Perception | 0.33 |
| Macaque IT cortex | Neural recordings | 0.84 |
| SlowFast | Video action recognition | 0.019 |
| ResNet-18 | Image classifier | 0.0018 |
| SimCLR ResNet-50 | Self-supervised image model | 0.0012 |
| VGG-16 | Image classifier | 0.00058 |
| ViT-L/32 | Vision transformer | 0.00032 |
| AlexNet | Image classifier | -0.00022 |
| ConvRNN | Recurrent video model | -0.0016 |
The image models in this table were tested with simulated adaptation, activity decay fitted to real IT neurons; the video models received drifting-stripe movies. Values near zero, positive or negative, mean no meaningful shift.
They know where things are
The models were not short of position information. When the team read position from their features in ordinary conditions, VGG-16 scored a correlation of 0.84 with true horizontal position and ResNet-18 scored 0.81, against about 0.60 for IT activity with more than 150 recorded units. The transformer ViT-L/32 was weaker at 0.59. AI vision networks can match or beat the brain on this basic task.
How accurately position can be read out in normal viewing (correlation with true horizontal position, from the paper)
But they do not adapt
Standard feedforward networks have no memory of the previous frame, so by design they give the same answer before and after adaptation. The team therefore built in “fatigue”: unit activity that decays over time, using time constants measured from the real IT neurons. Activity dropped as expected, but the position estimates still did not shift in the illusion’s direction.
Video models did not adapt either
The obvious next candidates were models built for motion. SlowFast, an action recognition network, and ConvRNN, which has recurrent connections and feedback, were shown adapter movies of drifting stripes followed by still images. Neither produced a meaningful shift. Temporal processing, as currently built into these AI vision models, was not enough.
Size of the illusion shift in degrees (bar length relative to macaque IT, 0.84 = 100%)
The neuralization test
The team then tried a shortcut. They measured how motion adaptation transformed IT activity and applied that same transformation to the features of VGG-16, ResNet-18 and SimCLR, a process they call “neuralization”. Once transformed, the models showed the same direction-opposite shifts as people and monkeys. The capacity is there; what is missing is a mechanism that produces the transformation by itself.
Why a Perceptual "Mistake" Can Be a Feature
It is tempting to read the result as AI vision beating human vision: the machine reports the true position, the person is fooled. The researchers see it differently.
Brains are not cameras
Biological vision does not just log pixel coordinates. It constantly adjusts to recent experience, which helps it stay sensitive to change and efficient with energy. The motion aftereffect is a side effect of that adjustment. Neuroscience News, summarising the paper, described the brain’s shifts as “an adaptive feature rather than a flaw”.
Getting the right answer is not enough
“If we want AI that works with humans and understands the world in more human-compatible ways, we cannot focus only on whether it gets the right answer,” Kar said in York’s announcement. “We also need to understand the computations that produce human perception and behavior.”
The NeuroAI argument
Kar described the study as capturing “the promise of NeuroAI”, using careful experiments to find computations that brains use and AI still lacks, then building them into better systems. On this view, illusions are not party tricks. They are probes that show which parts of perception a model has, and which it does not.
Why the AI Vision Gap Matters Outside the Lab
The experiment is narrow, but the question behind it is not. Whenever people and machines look at the same scene and need to agree, differences in how they see can matter.
Driving and robotics
A driver who has just watched traffic flow past, and a camera system that has not adapted at all, may judge a stationary object’s position slightly differently. In most cases the gap is tiny. In shared-control situations, where a person takes over from an automated system, understanding how each “sees” is part of designing a safe handover.
Medical imaging and inspection
Clinicians and inspectors scroll through moving or sequential images. If an AI vision assistant and a human reviewer are biased in different ways by recent images, they may disagree in ways that are hard to explain. Knowing that the model has no history effects, while the human does, is useful context for anyone designing those workflows.
Better benchmarks
The study’s authors say their results provide a new benchmark for models of dynamic vision. Most benchmarks reward models for matching ground truth. This one asks whether a model matches perception, which is a different and arguably harder test for computer vision.
Where motion-aware hardware fits
Other researchers are tackling time and motion from the hardware side. A team at KAIST recently built a stacked AI chip with fast and slow layers to recognise changes in motion over time. Approaches like that build history into the system, which is the ingredient the York study found missing.
What Teams Using AI Vision Should Take From This
Most organisations will never train a vision model from scratch. They buy one inside a camera system, an inspection tool or a document-processing service. The study still has practical lessons for them.
Test with sequences, not just single frames
Many evaluations score a model on still images. Real deployments see video and streams of images, where context builds up. If your use case involves motion, such as people moving through a site or products on a conveyor, test AI vision on realistic sequences and check whether recent frames change its answers.
Expect humans and machines to disagree differently
A human reviewer’s judgement is shaped by what they have just looked at. An AI vision model’s judgement, on this evidence, is not. Where both review the same material, build a process for resolving disagreements rather than assuming one side is always right.
Ask vendors about temporal behaviour
When buying an AI vision product, ask how it uses previous frames, whether it smooths or tracks objects over time, and how it was tested on video. Vendors that can answer clearly have usually thought about the problem.
Do not over-read one study
The York result is about a specific illusion and a specific set of research models. It does not mean commercial AI vision systems are unsafe. It means there is a measurable difference between how they and people process motion, and that difference is worth knowing about.
Try the Illusion Yourself
You do not need a laboratory to experience the effect the York team used to test AI vision. A minute and a moving scene are enough.
Find steady motion
A waterfall, a river, a scrolling credits sequence or a video of falling snow all work. The motion should be steady and fill a good part of your view.
Stare for about 30 seconds
Fix your eyes on one point in the middle of the motion and keep them still. The human participants in the study adapted for 30 seconds before their first test.
Look at something still
Switch your gaze to a still surface, such as a wall or a patterned rug. For a few seconds it will seem to drift the other way, or objects will seem slightly shifted. That is your visual system recalibrating, and it is exactly what the AI vision networks in the study did not do.
What the Study Does Not Show
The result is clear, but its scope is limited, and the paper is careful about that.
A small number of animals and one illusion
The brain recordings come from two monkeys, and the study uses one illusion in one direction of motion. The IT shift after leftward adaptation was not significant on its own. Broader tests, with other illusions and more subjects, would strengthen the conclusion.
Not the latest multimodal systems
The models tested are well-studied research architectures, including AlexNet, VGG-16, ResNet-18, ViT-L/32, SimCLR, SlowFast and ConvRNN. The paper does not test the large multimodal assistants that now answer questions about images. It would be wrong to assume they behave differently, but it would also be wrong to claim the study tested them.
IT is not the final step
The authors note that the shift in IT came with a drop in consistency that people did not show, which suggests areas downstream of IT add further processing before behaviour. IT is a likely place where perceptually aligned position is represented, not necessarily where the decision is made.
The fix was imposed from outside
Neuralization shows that existing AI vision features can express the illusion, but the transformation came from monkey data. Nobody has yet built a model that learns to produce it on its own.
What Comes Next for AI Vision Research
The paper ends with a research agenda for both neuroscience and AI vision engineering.
Mapping the brain circuit
The authors suggest recording from IT together with motion areas such as MT and V4, temporarily silencing those areas during adaptation, and recording across cortical layers to separate feedforward from feedback signals. That would show where the history effect comes from.
New model designs
The paper argues that the missing ingredient may be interaction between motion-selective and form-selective channels, combined with history-dependent gain control. That points towards architectures that explicitly couple form and motion pathways.
New training goals
Today’s AI vision models are mostly trained to label images or actions correctly. The authors suggest objectives that reward perceptual alignment, not just correct labels. That would be a significant change in how vision models are built and evaluated.
From lab to products
Turning a neuroscience finding into a better product is slow, as our piece on why translating AI advances to the real world is hard explains. Expect this to show up first in research benchmarks, then in specialised systems, long before it reaches consumer apps.
AI Vision and the Illusion: FAQ
What did the York study find?
People and monkey brain cells shifted their sense of an object’s position after watching motion, as the motion aftereffect predicts. The AI vision networks tested did not shift, even though they encoded position accurately.
What is the motion aftereffect?
It is an illusion in which, after watching steady motion, a still object seems to move or sit displaced in the opposite direction. The waterfall illusion is the classic example.
Which AI vision models were tested?
AlexNet, VGG-16, ResNet-18, ViT-L/32, SimCLR ResNet-50, and the video models SlowFast and ConvRNN.
Does this mean AI vision is worse than human vision?
Not exactly. The models were accurate about position. The difference is that they do not adapt to recent experience, which people and primates do. That matters most when humans and machines need to see the world the same way.
Where was the study published?
In Current Biology, by Elizaveta Yakubovskaya, Hamidreza Ramezanpour, Matteo Dunnhofer and Kohitij Kar. A preprint is on arXiv.
Could AI vision models learn this?
The study shows that their features can express the effect if an IT-like transformation is applied. Building a model that produces it on its own is the next challenge.
References and Further Reading
Visual Illusion Reveals AI Vision’s Blind Spot: York Study (York University release via Mirage News)
Optical Illusion Exposes Blind Spot in AI Vision (Neuroscience News)
Motion-Induced Object Position Adaptation in Macaque IT Cortex (Journal of Vision abstract, 2025)
Motion aftereffect (Wikipedia)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.