We have spent a combined forty years inside digital worlds. Building them, scaling them, watching them evolve from entertainment into economies, social systems, and deeply structured interactive environments. Twitch. Amazon. Game studios. Live video infrastructure. Production AI systems. Between the three of us, we have built, shipped, and operated across every layer of this stack.
And from inside those worlds, we noticed something the AI industry has been slow to see.
The environments that teach the most are game worlds, 3D simulations, interactive systems where actions have consequences and states are measurable. And they are almost entirely absent from the data pipelines powering today's frontier models. Instead, the default training substrate is the open web: scraped, noisy, unlicensed, and structurally biased toward whatever captures the most attention. That is what most AI learns from. The loudest slice of the internet.
We think that is a problem worth solving. And we think we are the right team to solve it.
The data problem is not a side issue
There is a growing body of research that says data quality is not secondary to model architecture. It is co-equal. The field has known this directionally since Chinchilla showed that many large language models were undertrained relative to their compute budgets.[1] But the more recent work has gone further. DataComp treated dataset design itself as the variable and showed that better curation can materially improve model quality under fixed compute.[2] FineWeb built a 15-trillion-token dataset from 96 Common Crawl snapshots and showed it could produce better open LLMs than other public pretraining sets, with its educational subset delivering particularly strong results on reasoning benchmarks.[3] DoReMi showed that optimizing training domain mixtures can reach baseline performance 2.6x faster.[4] And deduplication research showed that removing repetition from corpora reduces memorization and improves accuracy with fewer training steps.[5]
The message across all of this work is consistent: what you train on shapes what you get. Curation is not a nice-to-have. It is an algorithmic choice with first-order effects on capability and cost.
World models need worlds
This matters more now because of where the field is heading. AI is moving beyond text prediction toward spatial understanding, physical reasoning, and embodied action. The systems that will define the next generation, world models, robotics foundations, autonomous agents, all need training data that reflects how environments actually work. Not flat media. Not captions. Structured, interactive, consequence-rich data from systems where physics, causality, and state changes are real.
The research is already there. Genie showed that interactive environments can be learned from unlabeled video.[6] Genie 2 expanded this into controllable 3D world generation.[7] Genie 3 pushed further still, producing real-time interactive environments at 24 frames per second with visual consistency lasting minutes, not seconds.[8] V-JEPA 2 demonstrated that self-supervised learning on over a million hours of video can support understanding, prediction, and planning in the physical world.[9] NVIDIA's Cosmos framed the requirement directly: physical AI needs a digital twin of the policy model and a digital twin of the world, and building those twins requires petabytes of high-quality video data.[10]
That demand is not theoretical. Service robot sales reached nearly 200,000 units in 2024. In February 2026, World Labs raised over a billion dollars to build spatial intelligence.[13] Weeks earlier, Google DeepMind opened Project Genie to the public, putting real-time interactive world generation in the hands of users for the first time.[14] AMI Labs launched with a multi-billion-dollar valuation target to build world models from the ground up.[15] Every one of these efforts will be constrained by the same thing: the quality, structure, and provenance of their training data.
Our thesis
We call this space Artificial World Intelligence™. The idea is that the next generation of AI will be trained not on the open web, but on the world’s richest digital environments. Licensed, structured, and built from interactive systems where intelligence can learn cause and effect rather than just absorb the loudest signals on the internet.
Origin Lab exists to build the data infrastructure for that shift.
We acquire, create, and deliver rights-cleared, AI-enriched multimodal datasets from video games, 3D worlds, animation, and film. Every source is licensed before capture begins. Every dataset carries a full chain of custody, from source through capture, enrichment, and delivery. Our proprietary capture pipeline extracts not just video and audio, but player inputs, camera state, engine telemetry, and frame-level scene attributes across more than twenty queryable metadata categories.
This is not bulk media with tags. It is structured experience. The kind of data that teaches a model how environments behave, how actions produce consequences, and how states evolve over time.
Why we believe this matters
Training a text-based language model is computationally expensive. Training a video model is 10 to 100 times more expensive, because video is not a flat sequence of tokens. It is a spatiotemporal signal: every frame carries spatial information, every second carries temporal information, and the model must learn how those dimensions interact across time.[12] Meta used 6,144 H100 GPUs to train a single 30-billion-parameter video model. NVIDIA's Cosmos was trained on 9,000 trillion tokens from 20 million hours of data.[10] Generating a single 10-second video consumes GPU resources equivalent to thousands of text queries.[12]
World models go further still. They do not just generate pixels. They must learn physics, causality, and state changes. They must predict what happens next given an action, not just what looks plausible. That is a fundamentally harder problem, and the compute requirements scale accordingly. As these models grow in capability, the cost of training on bad data does not just increase. It multiplies. Every low-quality frame, every redundant sequence, every poorly structured clip wastes compute at a rate that makes text-era waste look modest by comparison.
That waste is not abstract. It is electricity: the IEA estimates global data centers consumed around 415 TWh in 2024 and projects roughly double that by 2030, with AI as the primary driver.[11] It is time: slower research cycles, longer training runs, more iterations to reach the same result. And it is a broken relationship with the people who built the environments this industry learns from. Game developers, 3D artists, filmmakers, and studios create the world's most complex and structured digital environments, and right now they are treated as a free extraction layer for someone else's model. Better data fixes all three problems. Fewer redundant training cycles, less wasted compute, and a system where the people who contribute to AI actually participate in the upside.[1][3][4][5] That is why every Origin Lab engagement includes clear licensing, usage tracking, and revenue sharing with rights holders.
What is ahead
We are not arguing against scale. We are arguing for better input, and for asking better questions. How do you score training utility across different environments? Which worlds teach causality rather than superficial correlation? How do licensed capture, engine telemetry, scene intelligence, and quality assurance become a real substrate for frontier research? As the field moves toward spatial reasoning, world simulation, and multimodal agency, these questions will matter more with every model generation.
Where this leads
And then there is the part that excites us most.
If you control the training data, you have a say in what AI learns. Not just what it sees, but what it practices. What behaviors it encounters most often. What patterns it internalizes as normal. In the worlds we work with, we can choose what the camera sees. We can watch characters cooperate, build, solve problems together, navigate complexity with care. We can create training environments where the default is not combat and extraction, but collaboration, resourcefulness, curiosity, and awe. We will have a lot more to say about this in a future post, because we think it is one of the most underexplored ideas in AI today.
Beyond curation, we see a longer road. The infrastructure we are building today, the capture pipelines, the enrichment and scoring systems, the rights-cleared data, can become a foundation for models built on data that was designed to be good from the start. Not scraped and cleaned after the fact. Designed, licensed, and structured with intention. We are already building our own proprietary capture software, and our plan is to keep expanding it and eventually put it in the hands of everyday people. We are a data company today. Where that leads is a story we are still writing.
The bigger vision is simple. One day, with the right provenance, licensing, security, and privacy protections in place, far more of the world's richest environments and experiences will responsibly contribute to how AI understands us. Not just game studios and film distributors. A mom walking her kids to school with a phone in her hand. An architect navigating a job site. A first responder moving through a building. Anyone whose lived experience of the world could help AI learn to be better, with their knowledge, their consent, and their share of the value.
That is the future we are working toward. And we started Origin Lab because we want to help build it.
Not because the world needed more scraped data. Because it needs better worlds to learn from.