Video game data is fast becoming one of the most sought-after raw materials in artificial intelligence, and your dodgy gaming skills are part of the supply. On 28 September 2026 WIRED reported that Worldmodeldata, a Cambridge startup advised by Yann LeCun, has licensed almost 1 million hours of gameplay from game studios. It plans to sell that footage, paired with every thumbstick twirl and trigger squeeze, to labs building “world models”: AI systems meant to understand how the physical world behaves.

The pitch is simple. Chatbots learned from an ocean of text, but robots, self-driving cars and drones need a different diet: pictures of a 3D space matched with the actions that changed it. Games record exactly that, at enormous scale, as a by-product of people playing them. Not everyone is convinced, and Nvidia’s world model lead thinks game physics is too crude for delicate robot work.

This article explains what WIRED reported, why world models are short of data, and what earlier research on Minecraft and Bleeding Edge shows. It compares the companies chasing video game data, checks the numbers behind the 1 million hours, and sets out the consent and data protection questions that come with recording how millions of people play.

What WIRED Reported About Video Game Data and Worldmodeldata

video game data ai world models gaming skills b block tree built from cubes

WIRED’s Joel Khalili opens with an image most players will recognise: “With the twirl of a thumbstick, squeeze of a trigger, and press of a few buttons, even an unskilled player can dance their way through a 3D video game environment.” Worldmodeldata is wagering that those clumsy sequences contain “a trove of information” for training new AI.

A thumbstick, a trigger and a trove

The article places the startup inside a wider argument. A “growing belief in corners of the AI industry” holds that LLMs, trained only on words, will be limited by their inability to navigate the physical world. They are “perhaps ill-equipped to pilot autonomous vehicles, steer robotic arms, or perform any other action that requires finesse and precision.”

That is why researchers including Fei-Fei Li and Yann LeCun have turned to world models. These systems need visual data coupled with action data, and the article notes that “there is no similar corpus of material available to world model labs” to match the text that trained chatbots.

A broker, not a model builder

Worldmodeldata does not build models itself. It packages controller inputs and other data that game studios already collect, which WIRED calls “an exhaust product available in massive quantities”, into training sets. The pitch is that it saves labs from striking individual agreements with dozens of studios.

“There are millions of great games, and they are more and more similar to the real world,” chief executive Rhea Loucas told WIRED. “Why don’t we take the vast, abundant, diverse experiences from video games, and teach AI?” She believes video game data will eventually make up most of the training material for world models.

Almost 1 million hours, sources unnamed

The company says it has licensed “almost 1 million hours’ worth of data” from studios behind popular games, although Loucas declined to name them. It also wants to create ways for individual players to be paid. Loucas claims the approach “could well lead to the GPT moment for world models”.

The original WIRED report balances that optimism with sceptical voices from Nvidia and the University of Surrey, covered below.

Why World Models Need Action Data, Not Just Words

video game data ai world models gaming skills c forklift carrying a crate

A world model tries to predict what happens next in an environment, especially after an agent does something. Where a chatbot predicts the next word, a world model predicts the next state of a scene: where the ball goes, whether the door opens, how the car ahead reacts.

Cause and consequence

“For world models, you need cause and consequence,” Xiatian Zhu, an associate professor specialising in AI at the University of Surrey, told WIRED. “On the internet, we have very little of this type of data.” Online video shows plenty of consequences, but it rarely records the actions that caused them.

That distinction matters. A clip of someone opening a fridge shows the door swinging, but not how hard the hand pulled. To learn to control a robot arm, a model needs to know both. WIRED gives the factory example: footage of the floor “coupled with information about how firmly an object should be gripped, with what torque it’s manipulated”.

A scaling bet

Researchers are “broadly working under the assumption” that world models will improve as their datasets grow, as LLMs did. The hypothesis “is yet to be fully tested”, WIRED notes, but if it holds, the shortage of suitable training data becomes one of the largest bottlenecks in the field. That bottleneck is the gap video game data is meant to fill.

Why the money is flowing now

The shortage has not deterred investors. AMI Labs, the company LeCun co-founded after leaving Meta, raised $1.03 billion in March 2026 at a $3.5 billion pre-money valuation. Fei-Fei Li’s World Labs raised $1 billion in February. Each dollar spent on compute for these labs needs data to chew on.

AMI’s chief executive Alexandre LeBrun predicted that “world models” would be the next buzzword, and that “in six months, every company will call itself a world model to raise funding.” Suppliers of the data are the obvious second wave.

How Video Game Data Captures Cause and Consequence

video game data ai world models gaming skills d apple on a round side table

A modern game logs, frame by frame, what the player saw and what the player pressed. That pairing, a picture of a 3D space plus the action taken in it, is the “cause and consequence” record Zhu says the internet lacks. Video game data also comes with the underlying 3D state of the scene, which ordinary video never has.

Frames paired with inputs

According to The Next Web’s report on the seed round, Worldmodeldata packages “the video, the player inputs and the underlying 3D state into clean datasets”. Tech.eu reported that the data comes from licensing agreements with game developers and communities, “including titles built on Unreal and Unity, rather than through web scraping.”

Corner cases at scale

Some labs generate their own data by attaching sensors to people and robots, but that yields small volumes and misses the odd situations a model will meet in the wild. “You can pay people to demonstrate pick-and-place tasks. But repetition alone won’t capture the disorder of the world,” says Nicole Fraenkel, a partner at Khosla Ventures.

“The corner cases are the ones to actually get right,” she adds. “The cost of error with a car, plane, drone, factory forklift, or autonomous quadruped is very high.” Games throw players into collisions, crowds, bad weather and chaos far more often than real life does, which is part of the appeal of video game data.

How the main data sources compare

Data sourceActions recorded?ScaleMain weakness
Online videoNo, must be inferredVastConsequences without causes
Sensor capture and teleoperationYes, preciselySmall and slowExpensive, repetitive, few corner cases
Physics simulation enginesYes, generatedAs large as compute allowsOnly as varied as the scenarios engineers write
Video game dataYes, controller inputsLarge, if studios license itGame physics takes shortcuts

The Research Record on Video Game Data: Minecraft and Bleeding Edge

video game data ai world models gaming skills e pair of dice

The idea that video game data can teach AI is not new, and two published projects show both its promise and its cost. Both are useful yardsticks for judging a claim of 1 million hours of video game data.

OpenAI’s Video PreTraining in Minecraft

In 2022 OpenAI published Video PreTraining (VPT). Searches for public Minecraft footage returned about 270,000 hours of video, which the team filtered to roughly 70,000 hours of “clean” play. That footage had no action labels, so contractors recorded 1,962 hours of play with their keyboard and mouse inputs logged.

A separate “inverse dynamics” model trained on that small labelled set learned to guess the missing actions, reaching 90.6% keypress accuracy. It then labelled the 70,000 hours. The resulting agent, fine-tuned with imitation learning and reinforcement learning, became the first computer agent to craft diamond tools, a task that takes skilled humans about 20 minutes, or 24,000 actions.

Microsoft’s Muse and seven years of Bleeding Edge

In February 2025 Microsoft Research and Ninja Theory published a World and Human Action Model in Nature, known as Muse. It was trained on Bleeding Edge, a four-versus-four online game, using 60,986 matches that yielded about 500,000 player trajectories. That came to more than seven years of gameplay and roughly 1.4 billion frames at 10 frames per second.

Crucially, that was genuine video game data, not inferred labels: every frame carried the real controller actions. Matches were only recorded when players had agreed to the game’s licence terms, and the paper describes the sessions as anonymised. The largest Muse model has 1.6 billion parameters.

What the scaling curves showed

VPT is the clearest public evidence that more gameplay produces new behaviour. Its paper reports that a model needed at least 10 hours of labelled data before any crafting appeared. Gains flattened after about 100 hours, and crafting tables only emerged once the foundation model saw over 5,000 hours.

OpenAI VPT: hours of Minecraft data at each milestone (log scale, figures from the VPT paper)
Labelled data before any crafting appears 10 hours
Labelled data where gains plateau 100 hours
Contractor play actually labelled 1,962 hours
Foundation data before crafting tables emerge 5,000 hours
Clean web video used for the full model 70,000 hours
Raw Minecraft video found by search 270,000 hours

The lesson for anyone buying video game data is that a small amount of accurately labelled play can unlock a much larger pool of footage, and that some behaviours only appear at scale.

Who Is Betting on Video Game Data

video game data ai world models gaming skills f stack of game cartridges

Worldmodeldata is one of several companies that see video game data as the route to physical AI. They differ mainly in whether they sell data or keep it.

Worldmodeldata: the Cambridge broker

The company emerged from stealth in July 2026 with a £7 million seed round led by London’s Iona Star Capital, according to Tech.eu. Lord Richard Allan, formerly Meta’s vice president of public policy, joined as chairman. Loucas frames the business as a neutral supplier that sells to every lab.

General Intuition: Medal’s in-house trove

General Intuition was spun out of Medal, a platform where players upload game clips. TechCrunch reported in June that it raised $320 million at a $2.3 billion valuation, bringing its disclosed funding to $454 million. Khosla Ventures led the round.

Medal’s “hundreds of millions of hours of uploaded gameplay” gave it a head start, but chief executive Pim de Witte says the key ingredient is the action labels embedded in those clips. “The human action data and reaction data you have in games is the key part to the emergence of intuition,” Vinod Khosla told TechCrunch. General Intuition keeps its video game data in-house to train its own models.

Origin Lab: publisher partnerships

San Francisco’s Origin Lab raised an $8 million seed round led by Lightspeed in May 2026. It says it has exclusive partnerships with more than 20 publishers covering over 50 titles, and a contract with a leading frontier lab. It is the closest direct rival to Worldmodeldata’s broker model.

Niantic Spatial: a map built by players

Niantic, the Pokémon Go developer, took a different route. Its spin-out Niantic Spatial has trained a model on 30 billion images captured by players in cities, and now helps Coco Robotics’ delivery robots find their way where GPS is weak. It is player-generated data, though from phone cameras rather than controllers.

CompanyModelDisclosed fundingData position
Worldmodeldata (Cambridge)Data broker£7m seed (July 2026)Almost 1m hours licensed from unnamed studios
Origin Lab (San Francisco)Data broker$8m seed (May 2026)20+ publishers, 50+ titles
General Intuition (New York)Builds its own agent$454m in totalMedal clips with action labels, kept in-house
Niantic SpatialGeospatial modelNot covered here30bn player images of urban landmarks
Nvidia CosmosOpen world foundation modelsCorporateReal video plus a custom physics engine

The Sceptics: Where Game Physics Falls Short

Not everyone believes video game data is the answer. The strongest objection comes from the company whose chips train most of these models.

Nvidia’s custom physics engine

Nvidia publishes a family of world models called Cosmos, launched at CES in January 2025. According to WIRED, Nvidia prefers “a custom engine it devised specifically to replicate real-world physics” as the backbone for its AI, rather than games. Nvidia also offers a pipeline it says can process and label 20 million hours of video in 14 days on its Blackwell platform.

Why manipulation is harder than navigation

Ming-Yu Liu, who leads world model development at Nvidia, told WIRED that models trained on video game inputs “are unlikely to fare well” at fine motor control. Game developers take shortcuts: a character dips a hand to collect an apple “without coding in the details”, such as the pressure each finger applies to stop it slipping.

“I would be more conservative on using video game data for manipulation,” Liu says. “The physics for manipulation is much more involved.” He thinks such data is better suited to world models that generate hyperrealistic video or 3D environments.

“Very coarse, approximate”

Zhu agrees. “Video games are, in essence, simulators. They do have some degree of physical grounding,” he says. “But they are very coarse, approximate.” Even supporters accept the point: Loucas expects video game data to be followed by fine-tuning on data from the specific real-world task.

TaskWhat the model must learnFit for game-trained models
Navigating a spaceWalls, doors, routes, other agentsStrong: games reward exactly this
Anticipating other road usersIntent, timing, rare eventsPromising, with real-world fine-tuning
Generating 3D scenes and videoVisual consistency over timeStrong, as Liu concedes
Grasping and manipulating objectsForce, friction, torque, slipWeak: game physics fakes these details

Checking the Numbers Behind 1 Million Hours of Video Game Data

Headline figures for video game data are hard to verify, because studios and buyers are rarely named. The public record still allows a few checks.

From target to licensed library

In July, Tech.eu reported that the seed money would help Worldmodeldata build “a library of one million hours of training data”, and TNW gave the deadline as the end of 2026. By 28 September the company told WIRED it had licensed almost that amount. That is fast progress, but “licensed” is not the same as cleaned, synchronised and delivered.

The “25 times” comparison

TNW reported that Worldmodeldata described 1 million hours as 25 times the largest existing dataset, a comparison TNW called “its own, and unverified”. Dividing 1 million by 25 implies a previous record of about 40,000 hours. VPT’s clean Minecraft set was roughly 70,000 hours, though its actions were inferred, and Muse’s seven-plus years of Bleeding Edge equal at least 61,320 hours (7 × 8,760).

On those published figures, the gap is closer to 14 to 16 times than 25. It is still a large dataset of video game data, and a stricter definition, such as counting only data with recorded controller inputs, could explain the difference.

Hours of gameplay in published and claimed datasets (bars scaled to 1 million hours; the smallest is widened to stay visible)
VPT contractor play with logged inputs 1,962 hours
Implied previous record (1,000,000 ÷ 25) 40,000 hours
Muse, 7+ years of Bleeding Edge (7 × 8,760) 61,320 hours
VPT clean Minecraft video, actions inferred 70,000 hours
Worldmodeldata, licensed (WIRED) ~1,000,000 hours

To put 1 million hours in human terms: a year has 8,760 hours, so the library equals about 114 years of non-stop play. General Intuition’s “hundreds of millions of hours” of Medal clips dwarf it, but that trove is not for sale.

A small team with big ambitions

TNW also reported that in July Worldmodeldata had “no finalised customer contracts and no revenue”, and about ten people including advisers and contractors. Its seed round had closed seven months before the announcement. Buyers should ask what has changed since then.

Every hour in these datasets was played by a person. That raises questions that text scraping made familiar and that video game data now inherits.

Licence terms and anonymised sessions

Microsoft’s Muse work offers a model of care. Its research blog says Bleeding Edge matches were recorded only if the player accepted the end-user licence agreement, and that the team worked with Microsoft’s compliance staff. The Nature paper describes the sessions as anonymised. A broker combining data from many studios inherits many different sets of terms.

Paying players for their play

Worldmodeldata says it wants to create ways for individual players to be compensated. General Intuition has launched Nerve, a marketplace where gamers can earn money, starting with data labelling and moving towards robot teleoperation. Both moves acknowledge that players are the ultimate source of the value.

UK data protection questions

Under UK GDPR, gameplay tied to an account, a voice chat or a distinctive play style may count as personal data. The ICO’s anonymisation guidance says data is only anonymous if people cannot reasonably be identified from it. Buyers of video game data should ask how identifiers, chat and usernames were stripped before licensing.

Lessons from earlier platform disputes

Streaming has already tested this ground. Amazon’s plan to use Twitch content for AI training prompted an opt-out row and later a lawsuit from streamers. Game studios licensing player data should expect similar scrutiny if players feel they were never asked.

What Video Game Data Means for Robotics and Autonomous Systems

If the bet on video game data pays off, the training recipe for physical AI starts to look like the one for chatbots: broad pre-training on cheap, plentiful data, then a small amount of expensive, task-specific data at the end.

Pre-training first, real-world data second

General Intuition’s demonstration hints at how much the second stage might shrink. TechCrunch reported that it took “just eight minutes of real-world robotics data” to fine-tune its model for a quadruped robot, and that data was gathered on the street rather than in the office where the robot was shown. One demonstration proves little, but the direction is clear.

Where businesses will feel it

Warehouse robots, delivery robots, drones and driver-assistance systems all need to cope with rare events. Cheaper pre-training could lower the cost of building such systems and speed up safety-critical AI testing. China is already writing data standards for embodied AI, a sign that regulators see training data for robots as strategic.

A new revenue line for game studios

For studios, video game data turns telemetry they already hold into a licensable asset. That may help smaller developers in particular, but it also creates obligations: clear consent, clean anonymisation and contracts that limit how buyers use the footage. General Intuition, for example, has said its agents will not be used to harm people.

How to Evaluate a Video Game Data Supplier

Organisations planning machine learning model development for robots or simulation should treat a video game data purchase like any other critical supply contract.

Provenance and licence scope

Ask which studios and titles supplied the footage, under what player terms, and whether the licence covers commercial training, fine-tuning and derived models. Unnamed sources may be commercially necessary, but the chain of consent still needs to be documented.

Action fidelity and synchronisation

Check whether actions in the video game data were recorded from real controller inputs or inferred afterwards, and how tightly they are synchronised with frames. Muse stored a timecode for every frame for exactly this reason. Poorly aligned inputs teach a model the wrong causes.

Physics realism and domain gap

Match the data to the task. Navigation and scene generation suit video game data far better than grasping, as Liu warns. Plan for a real-world fine-tuning stage and a test set that the game footage never touched.

Coverage of corner cases

Ask for statistics on the rare situations that matter to you, such as crowds, collisions, poor visibility or sudden hazards, rather than headline hours. A million hours of the same map is less useful than far fewer hours of varied play. Our AI and machine learning services team can help define those tests.

Video Game Data FAQs

What is video game data in AI training?

It is recorded gameplay: the frames a player saw paired with the controller or keyboard inputs they made, sometimes with the game’s 3D scene data. AI labs use it to teach world models how actions change an environment.

What is Worldmodeldata?

A Cambridge startup, founded by Rhea Loucas and advised by Yann LeCun, that licenses gameplay from studios and packages it for AI labs. It raised a £7 million seed round in July 2026 and told WIRED it has licensed almost 1 million hours.

What is a world model?

An AI system that predicts how an environment will change, especially in response to an action. World models are meant to help robots, autonomous vehicles and drones plan in the physical world, where LLMs trained on text struggle.

Can games really teach robots?

Partly. Research such as OpenAI’s VPT shows game play can teach complex behaviour inside a game, and companies report quick transfer to real robots after fine-tuning. Nvidia’s Ming-Yu Liu doubts it will work for delicate manipulation, because game physics takes shortcuts.

Do players get paid for their gameplay data?

Not usually today. Worldmodeldata says it wants to create ways to compensate individual players, and General Intuition runs a marketplace where gamers can earn money. Most data is licensed by studios under their existing player terms.

References and Further Reading