Research

    World models need
    world data.

    Origin Research is actively collaborating with AI researchers from Oxford and Google Research to provide training data, access, and support in driving breakthroughs that move Artificial World Intelligence™ forward faster and more efficiently. If you are interested in working with us, see our current tracks below.

    02/Details
    Research directions
    in depth.

    Each track includes a guiding research question and specific directions we're actively exploring with our partners.

    01
    Generative world-building pipelines

    How can we generate large, high-fidelity, non-redundant game worlds that are maximally useful for world model training, starting from existing engines/maps instead of a blank canvas?

    Using existing maps/engines as "seeds": what's the best way to learn structural priors (terrain types, building layouts, interaction density) and then extrapolate them 10 to 100x while avoiding copy-paste redundancy?

    Hybrid workflows: what's the right division of labor between generative systems and human curation/repair, especially for edge cases and physics glitches?

    Representation choices: should we model worlds at the level of pixels, geometry graphs, asset libraries, or some hybrid? How does that choice affect downstream utility (e.g., for robotics vs. pure vision)?

    Evaluation: how do we measure "world quality" for training (coverage, diversity, interaction richness, physics fidelity) vs. traditional game design metrics like "fun"?

    02
    Synthetic human players for capture + QA

    How can we build synthetic agents that navigate and interact with complex game worlds in a way that is both maximally informative for training and plausibly human-like?

    Behavior learning: given multimodal logs (video, inputs, webcam, mic), what's the best way to imitate or abstract human playstyles vs. optimizing purely for coverage/utility?

    Task conditioning: how do we specify and enforce arbitrary behaviors (e.g., "loop closures," "backwards traversal," "craft a house," static observation with camera panning) without hand-engineering per-game logic?

    Integration with existing tools: how far can we get by extending/controlling existing systems like Nvidia Nitrogen vs. building our own agent stack?

    Feedback loops: can we use world/value scoring and QA signals to adapt synthetic agent policies over time (e.g., avoid redundant regions or "low-value" behaviors)?

    03
    Scene intelligence for indexing and retrieval

    How can we build a scene understanding and indexing stack that is dramatically better than "generic VLM every 30 seconds" and tuned to the specific needs of world model and vision researchers?

    Model architecture: what specialized vision/attention models (or fine-tuned VLMs) are best suited for dense scene labeling across long videos (hours) with strict cost and latency constraints?

    Query design: how do we represent and index scenes so that complex semantic queries (200 hours of night-time city driving with rain, inside a car, with dialog) are cheap and accurate?

    Multi-source fusion: how do we combine video, in-engine telemetry, player inputs, and audio transcripts into a unified, searchable representation?

    Adaptivity: can we design a system that can be cheaply re-run or incrementally updated when new "hot" query dimensions emerge from the market?

    04
    World value scoring for training utility

    How can we systematically estimate the "training value" and fair economic price of a game or world using both static information and dynamic evidence?

    Feature design: which factors (physics fidelity, interaction diversity, traversal modalities, environment density, entity complexity) matter most for different downstream use cases, and how should scoring weights shift across them?

    Combining priors and evidence: how do we fuse engine metadata with empirical capture data into calibrated utility estimates with confidence intervals?

    Learning from comparables: can we build a "comps" framework that learns relative value from reference titles and generalizes to new or unreleased games across market tiers?

    Pricing surfaces: how should value-based pricing and quality tiers surface to licensors and buyers so they're explainable and persuasive?

    05
    Automated data quality and AI QA

    How can we automatically assess and incentivize high-quality human capture behavior (engaged, diverse, instruction-following) using the full multimodal signal we collect?

    Signal design: which signals (keyboard/mouse patterns, in-game trajectories, webcam, mic, eye-tracking, content features) are most predictive of "good" vs. "bad" capture sessions for different spec types (e.g., static observation vs. active exploration)?

    Real-time vs. offline: what parts of QA should run live (nudging, warnings, pause/stop) vs. as post-hoc scoring, and how do we keep both cheap and robust?

    Fairness & incentives: how do we avoid penalizing valid but unintuitive behaviors (e.g., long static pans) while still making it hard to game the system for maximum pay with minimum effort?

    Labeling and ground truth: how can we efficiently collect human judgments of capture quality to train and calibrate automated QA models at scale?

    03/Built for
    Worlds in motion.

    Every clip is a world unfolding - not a still frame.Physics, input, scene state, and camera - captured frame-by-frame inside the engine as the world reacts. The dynamic ground truth scraped video and pure simulation can’t deliver.

    World models

    Predict the next frame - and the next world. Train on engine-rendered video with ground-truth depth, semantics, and camera pose. The signals you can’t scrape.

    Embodied AI & robotics

    Break out of the box. Synthetic environments with real physics, full sensor stacks, and dense supervision - so policies transfer when the world stops being simulated.

    Autonomous vehicles

    More edge cases than you can drive. Rare conditions, long-tail interactions, and hazardous scenarios - captured with multimodal ground truth before they ever reach a real road.

    Multimodal foundation models

    Ten synchronized modalities per scene - pre and post-HUD RGB, depth, normals, audio, inputs, camera pose, physics state, and action labels. Pretrain on the structure scraped web video can’t carry.

    04/Open calls
    We’re funding
    work on…

    Concrete directions we want pushed. We provide data, compute credits, and engineering support for serious teams.

    • Causal world models that learn from interactive worlds, not scraped web video
    • Sim-to-real transfer where the “sim” already has physics, semantics, and ground-truth camera pose
    • Long-tail edge cases for AV perception - captured with multimodal ground truth before they happen on the road
    • Pretraining beyond RGB - depth, normals, camera pose, input, and engine state as first-class modalities
    • Provenance-aware training - model lineage, dataset rights, and reproducibility at scale
    Pitch your projectA paragraph is enough. We reply within a week.

    Want to collaborate?

    If you're a researcher working on world models, embodied intelligence, generative simulation, or data provenance - we'd love to hear from you.

    Get in touch