This is the fourth post in our series on Artificial World Intelligence™. In the first three, we explained why we started Origin Lab, defined what AWI means, and made the case for why rights-cleared data is technically and legally superior to the scraped alternative. This post is for a different audience. It is for the people who sign off on training budgets, approve compute purchases, and decide how much of their company's capital goes into AI infrastructure. If that is you, this post is a memo.
The short version: scraped and unstructured training data is the most costly line item in modern AI, and most companies are not counting it.
Compute is exploding. So is the waste
Anthropic CEO Dario Amodei has predicted that individual AI training runs could exceed $10 billion by 2026, and that AI companies will aim to build $100 billion training clusters by 2027.[1] Epoch AI's independent analysis supports the trajectory: the cost of training frontier AI models has grown 2.4x per year since 2016, and the largest runs are on pace to exceed one billion dollars by 2027.[2] OpenAI reportedly spent around $5 billion on research and development compute in 2024 alone, with the majority going to experiments and intermediate training runs rather than the final runs that produced released models.[3]
Those numbers are already staggering. But here is the part that matters for data quality: Epoch estimates that the ratio of total compute to final training run compute ranges from 1.2x to 4x, with a median of 2.2x.[2] That means for every dollar spent on the run that actually ships, another dollar or more goes to experiments, failed attempts, and iterations that did not work. A significant portion of that overhead is driven by data problems: duplicated content, domain mismatch, low-quality samples, provenance issues that surface late, and evaluation cycles that cannot diagnose where the failure came from because the data itself is poorly documented.
When training costs were in the low millions, sloppy data practices were an annoyance. When they are in the hundreds of millions, they are a capital allocation failure. When they hit ten billion, they are a board-level problem.
The evidence is clear
The research on data curation converges on one number: 20 to 40%. That is the range of compute savings that better data delivers for equivalent or better model performance, across multiple independent studies published in the last two years.[4][5][6][7] At a $100 million training budget, that is $20 million to $40 million saved. At a billion, $200 million to $400 million. Not a rounding error.
The pattern holds in fine-tuning too. Microsoft's phi-1 and Meta's LIMA both showed that small, carefully curated datasets can outperform much larger, unfiltered ones.[8][9] High-quality, task-shaped data often dominates brute-force quantity when compute, objectives, and model capacity are real constraints. Full details in the appendix.
And there is now a proof point at frontier scale. DeepSeek V3 achieved performance comparable to leading closed-source models for a total training cost of $5.6 million, while competitors spent $100 million or more on equivalent capability.[10] The difference was not a breakthrough in architecture alone. DeepSeek trained on 14.8 trillion carefully curated tokens, with the team explicitly citing a refined data pipeline designed to minimize redundancy while maintaining diversity. Their pre-training was, in the paper's own words, "remarkably stable," with zero irrecoverable loss spikes across the entire run.[10] When a company with constrained resources and export-restricted hardware matches the frontier by spending wisely on data, the argument that more compute always wins is no longer credible.
Anthropic's own president, Daniela Amodei, has made the same argument. Her governing principle for Anthropic's strategy is "do more with less." As she framed it to CNBC in January 2026, the next phase of AI competition will not be decided by who can afford the largest pre-training runs, but by who delivers the most capability per dollar of compute spent.[11]
What this means for world models
Everything above applies to language models. For world models, the stakes multiply.
As we wrote in our first post, training a video model is 10 to 100 times more expensive than training a text model.[12] World models go further: they must learn not just how scenes look but how environments behave, how physics works, how actions produce consequences over time. Every low-quality frame, every redundant sequence, every poorly structured clip wastes compute at a rate that makes text-era waste look modest.
NVIDIA's Cosmos was trained on 9,000 trillion tokens from 20 million hours of data.[13] Meta used 6,144 H100 GPUs for a single video model. The compute budgets for the next generation of world models will be larger still. In that environment, the difference between a well-curated, structured training corpus and a bulk scraped one is not a quality preference. It is a financial decision with direct consequences for how long training takes, how many iterations are needed, and how much of the resulting model's capability was actually earned from the data versus lost to noise.
The costs you are not counting
Training compute is the cost everyone talks about. But there are two other line items that belong in every AI CEO's budget, and most companies are pretending they do not exist.
The first is litigation. Anthropic's copyright settlement cost $1.5 billion, with attorney fees recently reduced from a requested $300 million to $187.5 million, pending final court approval in April 2026.[14] The settlement required destruction of the pirated training dataset and was limited to past use of training data, not outputs. Statutory damages under U.S. copyright law can reach $150,000 per infringed work, which means theoretical exposure for companies training on millions of unlicensed works extends into the hundreds of billions of dollars.[14] In January 2026, music publishers filed a separate $3 billion lawsuit against Anthropic.[15] Apple now faces three separate copyright lawsuits over Apple Intelligence, with authors alleging the company trained on pirated books from shadow libraries.[16] In March 2026, a single complaint named eight companies together: Anthropic, Google, OpenAI, Meta, xAI, Apple, Perplexity, and NVIDIA, alleging they all used pirated books to train their models.[17] There are now more than 70 active copyright infringement cases against AI companies in the United States alone, more than double the count at the end of 2024, and legal experts expect 2026 to be the peak year for new filings.[15] The Andersen v. Stability AI trial is set for September 2026.[18] This is not an abstract regulatory risk. It is a present, quantifiable financial exposure that scales with the size of your training corpus and the opacity of your sourcing.

The second is reputation. Look at how the public narrative is shifting around data sourcing. Companies that license their training data are building partnerships with creators, publishers, and IP holders. Companies that scrape are building legal dockets. The difference matters for hiring, because the best researchers and engineers increasingly want to work at companies that are building responsibly, not ones defending extraction. It matters for partnerships, because IP holders will work with companies they trust and sue the ones they do not. And it matters for the long game, because the AI companies that are seen as being on the right side of creators will have access to the people who actually understand how environments, interactions, and creativity work. World models are not a pure engineering problem. They are a problem that requires taste, domain knowledge, and deep familiarity with the worlds they are learning from. The companies that alienate the people who build those worlds are cutting themselves off from the expertise they need most.
Five questions for every AI CEO
If you oversee AI spend, here are five questions worth asking your team before the next expensive training run.
1. Can we trace every sample in our training corpus back to its source? Provenance is not just legal insurance. It is a debugging tool. When a model underperforms, teams that can trace failure to source quality, label quality, or domain mismatch fix the problem fast. Teams that cannot trace it run more experiments. More experiments cost more compute.
2. Is anyone actually auditing the data for quality and redundancy, or is the team pointing at volume? Scale without evaluation is how companies spend nine figures on a training run built on duplicated, noisy, contaminated content. Your data team should be able to tell you what fraction of the corpus is load-bearing for the target capability. If they cannot, you are buying uncertainty with compute.
3. Is there better data available that we are not buying because the upfront cost is higher? This is the false economy question. Licensed, structured, rights-cleared data costs more per unit than scraped data. Scraped data costs more per unit of model capability gained, because you need more of it, you retrain more often, and you carry legal and operational risk that never appears in the data line item.
4. Is procurement doing the right math on value per training dollar? Data procurement at most AI companies is still measured on cost per gigabyte or cost per token. That is the wrong metric at frontier scale. The right metric is cost per unit of capability earned from the data. By that measure, a high-signal licensed dataset that shaves weeks off a training run and improves benchmark performance is not an expense. It is a compute multiplier.
5. What is our refresh strategy? Data is not a one-time purchase. Models trained on stale corpora drift. The ability to refresh, update, and version training data is an operational capability, not a nice-to-have. If you do not know who to call for updated coverage six months from now, you do not have a data supply chain. You have a legacy problem.
The Origin Lab argument
We are not selling more data. We are selling better training economics.
Origin Lab provides licensed, structured, multimodal datasets from video games, 3D worlds, animation, and film. Every source is rights-cleared. Every dataset carries a full chain of custody. Every delivery includes engine telemetry, input capture, frame-level metadata, and structured scene annotations across more than twenty categories. And every engagement includes the ability to go back to the source for more, to request specific coverage, and to build refresh cycles into the pipeline.
For model trainers, that means cleaner corpora, fewer redundant training cycles, sharper domain fit, and less waste per unit of capability gained. For CEOs and CFOs, it means a data procurement strategy that functions as a compute multiplier rather than a cost center.
The companies that will lead the next era of AI will not be the ones that spent the most on compute. They will be the ones that got the most intelligence per dollar of compute spent. And that starts with what goes into the model, not how big the model is.
Appendix: The curation research in detail
For readers who want the specific findings cited above:
SoftDedup (He et al., 2024): reweighting training data by commonness reduces training steps by at least 26% while maintaining or improving model quality.[4]
D4 (Tirumala et al., 2023): careful selection on top of deduplication delivers roughly 20% efficiency gains at the 6.7 billion parameter scale.[5]
FineWeb (Penedo et al., 2024): a 15-trillion-token dataset built from 96 Common Crawl snapshots outperformed other public pretraining datasets, with its educational subset delivering particularly strong results on reasoning benchmarks.[6]
DataComp-LM (Li et al., 2024): model-based filtering enabled a 7 billion parameter baseline to reach approximately 64% on MMLU while using 40% less compute than the prior open-data state of the art.[7]
phi-1 (Gunasekar et al., 2023): a 1.3 billion parameter model trained on 6 billion tokens of carefully selected data outperformed much larger models on code benchmarks.[8]
LIMA (Zhou et al., 2023): a 65 billion parameter model fine-tuned on only 1,000 curated examples produced remarkably strong instruction-following behavior.[9]