Real-world AI is harder to build than the headlines suggest. Systems that score brilliantly in a lab often stumble in a dim living room, a busy hospital ward or a rainy street at night. In an essay published by The Conversation on 28 September 2026, University of Mississippi computer scientist Bo Wang sets out three reasons why: the generalisation gap, the sheer variety of human behaviour, and the resources needed to make models robust.

Wang’s examples come from computer vision, his own field, rather than from the language models behind chatbots. A camera could help a clinician assess a patient’s gait, or an assistive robot could notice an older person struggling to stand. Both only help if they keep working “in messy, unpredictable or low-visibility conditions”, which is exactly where lab-trained models tend to fail.

This article walks through each of the three challenges, using the numbers in Wang’s own research papers to show how large the gaps are. It explains what domain adaptation and vision-language models can and cannot fix, and closes with a practical readiness checklist for any organisation planning to put an AI system in front of real people.

What Bo Wang Argues About Real-World AI

real-world AI - real world ai lab to life translating ai advances b basketball in a ring stand

Wang, an assistant professor of computer and information science, studies “AI systems designed to understand human behavior—especially through visual cues like body movements and interactions with objects”. His essay was republished by Tech Xplore under a Creative Commons licence.

From self-driving cars to a patient’s gait

The piece opens by acknowledging remarkable recent progress, from self-driving cars to the LLMs behind chatbots. Yet “systems that perform well in controlled lab settings often fall short when deployed in homes, hospitals or on city streets.” The distance between those two settings is the subject of the essay.

Three gaps, one question of trust

Wang names “three main challenges to taking AI from the lab to the real world: generalization, human behavior and resources.” Each is a different kind of gap. The first is about conditions, the second about situations, the third about money and hardware. Together they decide whether real-world AI can be trusted.

Why his examples matter

Wang’s work sits in the most practical corner of AI research: vision systems meant to watch over people. Pose estimation for rehabilitation, fall detection at night and robots that help in the home are all cases where a failure is not a bad score on a leaderboard but a missed fall or an unsafe action. Leaderboards rank AI models on curated test sets; a care home does not. That makes his view of real-world AI unusually concrete.

Why Lab Success Does Not Guarantee Real-World AI

real world ai lab to life translating ai advances c candle in a holder

Most AI systems learn from curated datasets and are judged on benchmarks built from similar data. That is sensible for research, because it makes results comparable, but it hides how a model behaves when conditions drift.

Curated data, messy world

“Most AI vision systems are trained on well-lit, high-quality images, often from curated sources like motion capture studios or daytime recordings,” Wang writes. When the same systems meet “a dimly lit home, hospital room or nighttime street”, they struggle. “This is not because the technology is flawed, but because the data used to train it cannot fully represent the enormous variability of the real world.”

Benchmarks reward the average case

A headline accuracy figure averages over the test set. If the test set is mostly daylight images, a model can post a strong number while failing almost every night-time image. A test set drawn from the same sources as the training data will flatter any model. Real-world AI has to be judged on the difficult minority of cases, because those are often where it matters most.

Three challenges at a glance

ChallengeWhat goes wrongWang’s exampleResearch response
Generalisation gapFamiliar task, unfamiliar conditionsPose estimation in low lightUnsupervised domain adaptation (UDAPose)
Human behaviourNew combinations of actions and objectsRecognising unseen human-object interactionsFoundation and vision-language models (EZ-HOI, TokAG)
Resource gapRobust models need compute that labs and devices lackModels trained in data centres, deployed on small hardwareEfficient algorithms and open-source tools

Challenge 1: The Generalisation Gap in Real-World AI

real world ai lab to life translating ai advances d hospital bed with a pillow

The first challenge to real-world AI is the most familiar: a model trained in one set of conditions is asked to work in another.

What pose estimation does

Human pose estimation detects key points on the body, such as joints, and joins them into what Wang calls “a digital skeleton”. It has been studied for health and rehabilitation applications, including gait measurement, monitoring rehabilitation exercises and spotting movement patterns linked to the risk of falling. A 2021 review by Johns Hopkins researchers noted it could let a clinician run a motor assessment in a patient’s home using a phone.

Why low light breaks models

“In well-lit environments, modern AI models do this with impressive accuracy. But in low or uneven lighting, performance can drop sharply,” Wang writes. The model “has learned from clean, consistent data and isn’t prepared to generalize to tougher conditions.”

The size of that drop is striking. In Wang’s own CVPR 2026 paper, a pose model trained only on well-lit images scored 60.1 average precision (AP) on well-lit test images. On the “hard” low-light split of the ExLPose benchmark it scored 0.4, and on the “extreme” split 0.2. That is not a gentle decline; it is a collapse that no lab accuracy figure would reveal about real-world AI.

Pose estimation accuracy by lighting (AP out of 100, from the UDAPose paper’s ablation table)
Well-lit images, model trained on well-lit only 60.1 AP
Normal low light, model trained on well-lit only 3.4 AP
Hard low light, model trained on well-lit only 0.4 AP
Well-lit images, UDAPose 67.3 AP
Normal low light, UDAPose 38.7 AP
Hard low light, UDAPose 28.0 AP

Why more night-time data is not a simple fix

“Would collecting more nighttime training data solve the problem?” Wang asks. It would help, but “when a person’s joints are hard to see in an image, they are also hard for a human annotator to label accurately.” Low light can obscure information rather than just darken it, and night scenes vary with cameras, shadows, glare and noise. Labelling examples that cover all of that “is therefore difficult and expensive.”

How UDAPose Narrows the Low-Light Gap

real world ai lab to life translating ai advances e traffic cones with stripes

Wang’s group responded with UDAPose, accepted at CVPR 2026 and led by Haopeng Chen with co-authors including Robby T. Tan and Yixin Chen.

Unsupervised domain adaptation

The method uses unsupervised domain adaptation: a model trained on well-lit, labelled images is adapted to low light “without needing manual body-joint labels for the low-light images”. The aim is to bridge the “domain gap” between clean training data and the conditions where real-world AI is used.

Synthesising realistic darkness

Earlier approaches darkened well-lit images by hand-crafted rules, which “oversimplify noise patterns”, or with learned translators that lose fine detail. UDAPose adds two modules that inject high-frequency detail taken from real low-light images, so its synthetic night scenes look more like genuine ones.

Balancing what it sees with what it knows

A third module, “Dynamic Control of Attention”, lets the model lean less on unreliable pixels and more on learned knowledge of how human bodies are built. Wang describes it as balancing “uncertain visual evidence with its learned knowledge of human body structure.”

How the results compare

Method (trained without low-light labels)Well-lit APHard low-light APExtreme low-light AP
RFormer (image enhancement)60.00.30.8
QuadPrior (image enhancement)60.24.60.3
EnCo (domain adaptation)60.016.22.9
ELLA (domain adaptation)61.517.23.4
CycleGAN (image translation)61.317.93.3
UDAPose67.328.011.7

On the hard split, UDAPose’s 28.0 AP is 10.1 points above the best previous method, a 56.4% relative gain. In a cross-dataset test on EHPT-XC, a separate set of RGB and event-camera recordings, it gained 7.4 AP, or 31.4%. Simply brightening images before detection, the approach of the enhancement methods, barely helped.

More synthetic data, smaller returns

The paper also varied how many synthetic low-light images were used. Returns shrink as the pile grows, a pattern that matters for any team budgeting real-world AI data collection.

Hard low-light AP by number of synthetic training images (UDAPose, bars scaled to 30 AP)
4,000 images 22.4 AP
8,000 images 24.8 AP
12,000 images 26.2 AP
16,000 images 27.3 AP
20,000 images 28.0 AP

The first 4,000 extra images (4,000 to 8,000) added 2.4 AP; the last 4,000 (16,000 to 20,000) added 0.7.

Far from solved

Wang is careful: the method “substantially improved performance on two low-light benchmarks, although the broader challenge is far from solved.” Even the best result, 11.7 AP on the extreme split, would be unusable for a fall detector. Real-world AI in the dark still needs better sensors, better data and honest testing.

Challenge 2: Human Behaviour Is the Hardest Real-World AI Test

real world ai lab to life translating ai advances f cutting board with a tomato and knife

The second gap facing real-world AI is about situations rather than conditions. People do an astonishing range of things with the objects around them.

Human-object interaction detection

In computer vision, recognising “actions like cutting a tomato or passing a basketball” is called human-object interaction (HOI) detection. It needs “not just object detection, but an understanding of context and intent,” Wang writes.

A chair can be sat on, carried, dragged or stacked

“The real challenge is scale,” he argues. “You can sit on a chair, carry it, drag it or stack it. Each is a distinct interaction.” A cup “might be used for drinking, washing, pouring or handing something to another person”, and an assistive robot would need to tell these apart. “It’s simply not feasible to collect and label data for every possible combination.”

What a benchmark covers

The widely used HICO-DET benchmark contains 600 interaction categories spread over 80 object categories, an average of 7.5 per object, across 47,776 images. That sounds large until you picture a real kitchen, workshop or care home, where the list of things people do with objects never ends. Real-world AI will always meet combinations nobody labelled.

Teaching Models Interactions They Have Never Seen

Wang’s lab studies “how an AI system can detect interactions it has never seen during training”, which researchers call a generalisation-to-unseen-classes problem.

Zero-shot HOI detection with EZ-HOI

In EZ-HOI, presented at NeurIPS 2024, the team adapted vision-language models with prompt learning, guided by descriptions from LLMs, so they could detect unseen interaction classes. The paper reports state-of-the-art zero-shot results with only 10.35% to 33.95% of the trainable parameters of existing methods.

Affordance grounding with vision-language models

A more recent study, TokAG, accepted to ECCV 2026, uses large vision-language models “to identify the parts of an object that support particular actions” without task-specific labelled training. It picks the output token whose attention focuses on the object, turning it into a heat map. It improved a standard localisation score by 10.7% on the unseen split of the AGD20K benchmark and by 29.7% on HICO-IIF.

Two generalisation problems, two solutions

Wang draws a useful distinction. In low-light pose estimation, “the system is trying to recognize a familiar type of movement under unfamiliar environmental conditions.” In HOI detection, “the action-and-object combination itself may be new.” Both are generalisation problems, “but they require different solutions”. Teams building real-world AI should identify which kind of novelty their system will face.

Challenge 3: The Resource Gap Behind Real-World AI

Even when researchers know how to make real-world AI more robust, they may not be able to afford it.

Industry built over 90% of notable models

“Today’s most powerful AI systems depend on massive computational resources,” Wang notes, often beyond university labs and public-sector researchers. According to the 2026 AI Index Report, industry produced over 90% of notable AI models in 2025. The same chapter notes that the most capable models no longer disclose training code, parameter counts or dataset sizes.

From data centre to device

“A model developed using large computing clusters may ultimately need to operate in a clinic, school, robot or home device, which means running on much smaller hardware,” Wang writes. Such systems may also need to respond quickly “without continuously sending sensitive data to the cloud.” That mismatch between training and deployment is where a lot of real-world AI quietly breaks.

Small models can keep up

There is encouraging evidence that size is not everything. In the TokAG paper, the 2-billion-parameter version scored 1.514 on the unseen AGD20K localisation measure against 1.549 for the 32-billion version, a gap of about 2% for a model 16 times smaller. The AI Index also highlights OLMo 3.1 Think 32B, which achieves results comparable to Grok 4 on several benchmarks with nearly 90 times fewer parameters, through pruning, deduplication and curation.

Wang’s conclusion is that bridging the gap “requires not just better models, but more efficient algorithms and open-source tools that broaden access to advanced AI.” His group publishes its code, as UDAPose’s public repository shows.

What it means for smaller organisations

The same gap applies to businesses. Few firms can train a frontier model, and most do not need to. Adapting an open model to their own cameras, lighting and tasks, as UDAPose adapts a pose model to the dark, is often a better route to real-world AI than chasing the largest system. Freely available AI tools and pretrained models make that practical. The skills involved are data collection, careful evaluation and deployment engineering.

Why Real-World AI Failures Are Failures of Trust

“These challenges aren’t just academic,” Wang writes. “They affect how and whether AI tools can be trusted in the real world.”

A fall detector that fails at night

He asks readers to imagine “a fall-detection system that fails at night, a workplace safety monitor that misses a crucial action or a robot that misunderstands what a person is doing”. These “aren’t just minor bugs—they’re breakdowns in trust, usability and sometimes even safety.” The chart above shows how easily the first could happen.

Everyday environments, amplified risks

AI systems are now deployed “in homes, hospitals, factories, vehicles and other everyday environments”. If they cannot adapt, “their usefulness becomes limited and their risks amplified.” Builders of safety-critical AI, from drones to autonomous trucks, make the same argument: the rare case is the one that counts.

Robustness as a design goal

That is why, Wang says, researchers are “increasingly focusing not just on accuracy in ideal conditions, but on robustness, generalization and real-world readiness.” For buyers of real-world AI, robustness should be a written requirement, not an assumption.

Designing for graceful failure

A system that knows when it is unsure can be safer than one that is simply more accurate. Confidence scores, alerts when lighting drops below the tested range and a clear hand-off to a person turn silent failures into visible ones. Wang’s essay does not cover this layer, but it is where many deployments succeed or fail.

A Real-World AI Readiness Checklist for Organisations

The essay is written for a general audience, but its lessons for real-world AI translate directly into procurement and project practice. Any organisation deploying computer vision or other perception systems can apply them.

Test where the system will live

Collect a test set from the actual site, camera and lighting, including night, glare and clutter. Do this before signing a contract. Any AI strategy for perception systems should start from that site survey.

Measure the tail, not the average

Ask real-world AI vendors for accuracy on the hardest conditions separately, as the ExLPose splits do. A strong average can hide a near-zero score where it matters.

Plan for situations nobody labelled

List the interactions and objects the system may meet, then ask how it handles ones outside its training data. Zero-shot methods help, but a human fallback is still needed.

Budget for edge hardware and privacy

Decide early whether the model must run on the device. On-device inference protects sensitive video and cuts delay, but it constrains model size. Teams planning machine learning model development should prototype on the target hardware from the start.

Question to askWhy it mattersWarning sign
Was it tested on our site’s lighting and cameras?Low light can cut pose accuracy from 60 AP to near zeroOnly daylight or studio test data
What is accuracy on the hardest 10% of cases?Averages hide failures in rare conditionsA single headline figure
How does it handle unseen objects or actions?Labelled data never covers every combinationNo stated behaviour for unknown inputs
Where does inference run?Cloud inference sends sensitive data off siteAccuracy quoted only for a data-centre model
What happens when it is unsure?Trust depends on graceful failureNo confidence score or human hand-off

Real-World AI FAQs

Why does AI fail outside the lab?

Mainly because training data cannot capture the full variety of real conditions and behaviours. Models learn from clean, curated examples and then meet darkness, clutter, new objects and unusual actions they were never shown.

What is the generalisation gap?

It is the difference between how a model performs on data like its training set and how it performs on new conditions. In Wang’s low-light research, a pose model that scored 60.1 AP in good light scored 0.4 AP in hard low light.

What is unsupervised domain adaptation?

A technique for adapting a model trained on labelled data from one setting, such as daylight, to another setting, such as night, without labelling the new data by hand. UDAPose does this by synthesising realistic low-light images.

Can vision-language models recognise actions they were not trained on?

Partly. Methods such as EZ-HOI and TokAG use the broad knowledge in vision-language models to detect interactions and object parts without task-specific labels, with measurable gains on unseen test splits. They reduce but do not remove the need for testing.

How can a business make real-world AI more reliable?

Test on data from the actual deployment site, report accuracy on the hardest conditions separately, plan for unseen situations with a human fallback, and choose hardware early so the model can run where it is needed.

References and Further Reading