Origin Research is actively collaborating with AI researchers from Oxford and Google Research to provide training data, access, and support in driving breakthroughs that move Artificial World Intelligence™ forward faster and more efficiently. If you are interested in working with us, see our current tracks below.
Each track is a collaboration between Origin Lab researchers, university partners, and frontier AI labs - with access to our full licensed dataset.
Procedural and learned approaches to synthesize high-resolution 3D environments that approximate the observable world.
Autonomous agents that navigate, interact, and generate training-quality capture sessions at scale.
Vision models that decompose scenes into searchable, structured representations for downstream use.
Proprietary scoring engine evaluating training utility across traversal density, interaction richness, physics fidelity, and combinatorial trajectory diversity. Determines value and pricing before capture begins.
AI-driven QA pipelines that monitor capture quality in real time, detect artifacts, reduce noise, and ensure dataset integrity at scale.
Each track includes a guiding research question and specific directions we're actively exploring with our partners.
How can we generate large, high-fidelity, non-redundant game worlds that are maximally useful for world model training, starting from existing engines/maps instead of a blank canvas?
Using existing maps/engines as "seeds": what's the best way to learn structural priors (terrain types, building layouts, interaction density) and then extrapolate them 10 to 100x while avoiding copy-paste redundancy?
Hybrid workflows: what's the right division of labor between generative systems and human curation/repair, especially for edge cases and physics glitches?
Representation choices: should we model worlds at the level of pixels, geometry graphs, asset libraries, or some hybrid? How does that choice affect downstream utility (e.g., for robotics vs. pure vision)?
Evaluation: how do we measure "world quality" for training (coverage, diversity, interaction richness, physics fidelity) vs. traditional game design metrics like "fun"?
How can we build synthetic agents that navigate and interact with complex game worlds in a way that is both maximally informative for training and plausibly human-like?
Behavior learning: given multimodal logs (video, inputs, webcam, mic), what's the best way to imitate or abstract human playstyles vs. optimizing purely for coverage/utility?
Task conditioning: how do we specify and enforce arbitrary behaviors (e.g., "loop closures," "backwards traversal," "craft a house," static observation with camera panning) without hand-engineering per-game logic?
Integration with existing tools: how far can we get by extending/controlling existing systems like Nvidia Nitrogen vs. building our own agent stack?
Feedback loops: can we use world/value scoring and QA signals to adapt synthetic agent policies over time (e.g., avoid redundant regions or "low-value" behaviors)?
How can we build a scene understanding and indexing stack that is dramatically better than "generic VLM every 30 seconds" and tuned to the specific needs of world model and vision researchers?
Model architecture: what specialized vision/attention models (or fine-tuned VLMs) are best suited for dense scene labeling across long videos (hours) with strict cost and latency constraints?
Query design: how do we represent and index scenes so that complex semantic queries (200 hours of night-time city driving with rain, inside a car, with dialog) are cheap and accurate?
Multi-source fusion: how do we combine video, in-engine telemetry, player inputs, and audio transcripts into a unified, searchable representation?
Adaptivity: can we design a system that can be cheaply re-run or incrementally updated when new "hot" query dimensions emerge from the market?
How can we systematically estimate the "training value" and fair economic price of a game or world using both static information and dynamic evidence?
Feature design: which factors (physics fidelity, interaction diversity, traversal modalities, environment density, entity complexity) matter most for different downstream use cases, and how should scoring weights shift across them?
Combining priors and evidence: how do we fuse engine metadata with empirical capture data into calibrated utility estimates with confidence intervals?
Learning from comparables: can we build a "comps" framework that learns relative value from reference titles and generalizes to new or unreleased games across market tiers?
Pricing surfaces: how should value-based pricing and quality tiers surface to licensors and buyers so they're explainable and persuasive?
How can we automatically assess and incentivize high-quality human capture behavior (engaged, diverse, instruction-following) using the full multimodal signal we collect?
Signal design: which signals (keyboard/mouse patterns, in-game trajectories, webcam, mic, eye-tracking, content features) are most predictive of "good" vs. "bad" capture sessions for different spec types (e.g., static observation vs. active exploration)?
Real-time vs. offline: what parts of QA should run live (nudging, warnings, pause/stop) vs. as post-hoc scoring, and how do we keep both cheap and robust?
Fairness & incentives: how do we avoid penalizing valid but unintuitive behaviors (e.g., long static pans) while still making it hard to game the system for maximum pay with minimum effort?
Labeling and ground truth: how can we efficiently collect human judgments of capture quality to train and calibrate automated QA models at scale?
Every clip is a world unfolding - not a still frame.Physics, input, scene state, and camera - captured frame-by-frame inside the engine as the world reacts. The dynamic ground truth scraped video and pure simulation can’t deliver.
Predict the next frame - and the next world. Train on engine-rendered video with ground-truth depth, semantics, and camera pose. The signals you can’t scrape.
Break out of the box. Synthetic environments with real physics, full sensor stacks, and dense supervision - so policies transfer when the world stops being simulated.
More edge cases than you can drive. Rare conditions, long-tail interactions, and hazardous scenarios - captured with multimodal ground truth before they ever reach a real road.
Ten synchronized modalities per scene - pre and post-HUD RGB, depth, normals, audio, inputs, camera pose, physics state, and action labels. Pretrain on the structure scraped web video can’t carry.
Concrete directions we want pushed. We provide data, compute credits, and engineering support for serious teams.
If you're a researcher working on world models, embodied intelligence, generative simulation, or data provenance - we'd love to hear from you.