Origin Research is actively collaborating with AI researchers from Oxford and Google Research to provide training data, access, and support in driving breakthroughs that move Artificial World Intelligence™ forward faster and more efficiently. If you are interested in working with us, see our current tracks below.
Each track is a collaboration between Origin Lab researchers, university partners, and frontier AI labs - with access to our full licensed dataset.
Fine-tuning foundation vision models into single-step depth and surface-normal predictors using engine-captured ground truth instead of pseudo-labels.
Training dense segmentation models on engine-emitted instance and semantic masks rather than human or model-generated annotations.
Reconstructing dynamic 3D scenes over time from captured RGB, depth, and ground-truth camera trajectories.
Systematically estimating the training value and fair economic price of a game world from statistical relevancy and diversity.
A scene understanding and indexing stack that makes a massive gameplay corpus semantically searchable.
Automatically assessing and incentivizing high-quality human capture, from real-time signals to corpus-level audits.
Building human-like synthetic agents that play games to generate maximally informative training capture at scale.
Procedural and learned synthesis of large, high-fidelity, non-redundant game worlds purpose-built for training.
Training interactive world models on engine capture, starting with a single FPS and scaling to a multi-game and multiplayer system.
Publishing an open benchmark that scores world models on action fidelity, controllability, and physical consistency against engine-measured ground truth.
Publishing an open benchmark that measures the training value of synthetic data against real engine capture of equal size.
Each track includes a guiding research question and specific directions we're actively exploring with our partners.
Fine-tuning foundation vision models into single-step depth and surface-normal predictors using engine-captured ground truth instead of pseudo-labels.
Every Origin Lab session ships pixel-perfect depth straight from the game engine, giving us supervision that inference-based pipelines can only approximate. We fine-tune diffusion backbones into single-step depth predictors conditioned on RGB, then extend the same recipe to surface normals captured at the render pass. We benchmark against state-of-the-art monocular estimators across depth edges, far field, fast camera motion, temporal flicker, and metric scale.
Training dense segmentation models on engine-emitted instance and semantic masks rather than human or model-generated annotations.
Game engines know what every pixel is before it is rendered. We capture object identity, material, and class masks directly at the render pass, producing dense, temporally consistent segmentation at a scale and accuracy no annotation pipeline can match. These streams support dynamic-object and boundary research while conditioning world models, scene indexing, and controllable generation.
Reconstructing dynamic 3D scenes over time from captured RGB, depth, and ground-truth camera trajectories.
Combining per-frame depth, camera intrinsics, and measured pose telemetry, we reconstruct geometry that moves, deforms, and interacts over time. This lets us isolate reconstruction quality from pose error and compare measured versus estimated signals across fast motion and large environments.
Systematically estimating the training value and fair economic price of a game world from statistical relevancy and diversity.
Not all captured hours are equal. We estimate how much a world contributes to a target corpus by measuring diversity, redundancy, and task coverage. Structural priors combine with evidence from early capture to produce comparative rankings, explainable pricing surfaces, and capture priorities.
A scene understanding and indexing stack that makes a massive gameplay corpus semantically searchable.
We build specialized models for dense scene labeling and fuse their output with telemetry, input, and audio. A semantic query layer can retrieve precise combinations of environments, events, mechanics, and conditions across titles, while adaptive re-indexing keeps the archive current as models improve.
Automatically assessing and incentivizing high-quality human capture, from real-time signals to corpus-level audits.
We research predictive signals for frame pacing, duplicate frames, sync drift, and coverage behavior, comparing live guidance with post-hoc scoring. Automated pipelines evaluate fitness for each downstream task while a parallel thread studies fair incentives and honest ground truth for the QA models themselves.
Building human-like synthetic agents that play games to generate maximally informative training capture at scale.
Agents learn from multimodal gameplay logs and accept task conditioning for directed behaviors such as exploration, mechanic coverage, and interaction stress tests. A feedback loop scores their footage and adapts policies toward the capture the corpus is missing.
Procedural and learned synthesis of large, high-fidelity, non-redundant game worlds purpose-built for training.
We explore structural priors from existing maps, human-guided synthesis, and the choice between generating pixels or engine-ready geometry. Evaluation is anchored to downstream training utility: whether generated worlds measurably improve models and supply the diversity the corpus lacks.
Training interactive world models on engine capture, starting with a single FPS and scaling to a multi-game and multiplayer system.
A frozen visual codec builds a latent cache over synchronized RGB, depth, and player input; an action-conditioned dynamics model learns on top. The program scales from one first-person title to multi-game transfer and multiplayer interaction, evaluated on commanded-action and camera-trajectory recoverability.
Publishing an open benchmark that scores world models on action fidelity, controllability, and physical consistency against engine-measured ground truth.
Held-out capture with true depth, poses, and action logs powers standardized probes for action fidelity, camera control, and physical consistency. Open leaderboards and baselines make results comparable across labs and establish a reference yardstick beyond visually plausible video.
Publishing an open benchmark that measures the training value of synthetic data against real engine capture of equal size.
Fixed downstream tasks and real held-out evaluation sets ask how much a generated corpus improves a model relative to equally sized real capture. The protocol exposes where synthetic data helps, merely pads, or actively harms, with baselines and tooling any lab can use.
Every clip is a world unfolding - not a still frame.Physics, input, scene state, and camera - captured frame-by-frame inside the engine as the world reacts. The dynamic ground truth scraped video and pure simulation can’t deliver.
Predict the next frame - and the next world. Train on engine-rendered video with ground-truth depth, semantics, and camera pose. The signals you can’t scrape.
Break out of the box. Synthetic environments with real physics, full sensor stacks, and dense supervision - so policies transfer when the world stops being simulated.
More edge cases than you can drive. Rare conditions, long-tail interactions, and hazardous scenarios - captured with multimodal ground truth before they ever reach a real road.
Ten synchronized modalities per scene - pre and post-HUD RGB, depth, normals, audio, inputs, camera pose, physics state, and action labels. Pretrain on the structure scraped web video can’t carry.
Concrete directions we want pushed. We provide data, compute credits, and engineering support for serious teams.
If you're a researcher working on world models, embodied intelligence, generative simulation, or data provenance - we'd love to hear from you.