Our Data

    A library of worlds,
    structured for training.

    Twenty-plus publisher partners. Fifty-plus licensed titles. Every recording carries 10 modalities frame-locked on one clock - RGB before and after the HUD, depth, normals, audio, inputs, camera, physics, and action labels - delivered through a single API, with more on the way.

    See the API
    Open datasets

    Research releases, open to explore.

    Explore the data, methods, and checkpoints behind our latest research. The complete release record lives on Hugging Face.

    Origin Lab on Hugging Face
    01 - Dataset

    Game Recordings v3: synchronized multimodal gameplay

    Human gameplay captured in-engine at 1080p and 60 FPS, with RGB, depth, normals, audio, actions, camera, and telemetry aligned on one shared frame clock.

    144sessions145.1hours10games4.9 TBsize
    Synchronized multimodal gameplay recordings from four games
    02 - Dataset

    Game-Depth: dense engine depth from licensed games

    A depth model trained only on video game frames beats purpose-built synthetic data on real KITTI photographs, with 4x less training data.

    48,615frames1080pRGBDensedepth101 GBsize
    Game-Depth predictions compared with ground truth and public depth models on KITTI driving scenes
    03 - Dataset

    Frame-Synced Multiplayer: eight players, one frame clock

    Eight PCs. Eight players. One frame clock. Every prior multi-view gameplay dataset was rendered from a replay; this one was captured live on independent machines.

    16matches108player-hours3.45Mshared frames3.8 TBsize
    Eight live player views synchronized on one frame clock
    04 - Model

    Lotus Game-Depth checkpoints

    Depth model checkpoints

    Released depth checkpoints make the Game-Depth result directly inspectable and reproducible.

    Lotus Game-Depth model repository preview

    Every release ships under the Origin Lab Data License with a 90-day internal evaluation track for labs. Commercial licensing requires a direct agreement. How release terms apply

    01/Live capture
    See it on real gameplay.

    Real licensed gameplay, captured in-engine at a 60 fps CFR standard. Events, Camera, Input, and Feed are different views into one synchronized record: ten modalities, game state, and in-engine action events on the same timeline.

    02/What's captured
    Time-aligned,
    frame by frame.

    What you just watched is 10 modalities captured at the engine level - depth, normals, telemetry, inputs, and action labels locked to the same clock as the rendered frame, at the source. They cannot be inferred from video alone.

    Pre-HUD RGB

    Pristine gameplay at up to 4K/60fps. HUD, menus, and UI stripped at the engine level before recording - your model trains on the raw 3D world, not on screen capture.

    Post-HUD RGB

    The same frame exactly as the player saw it, HUD and interface included. For agents and models that need to read the screen, not just the world.

    Depth maps

    Per-pixel depth straight from the renderer, not estimated after the fact. The same maps behind the slider in the demo above.

    Surface normals

    Per-pixel surface orientation from the engine. Geometry, lighting, and material structure for reconstruction and 3D understanding.

    Audio

    In-game audio captured as separate tracks. Dialogue, environmental sound, and effects at full fidelity.

    Keyboard inputs

    Exact key sequences straight from hardware, with press and release timing. Not inferred from video.

    Mouse inputs

    Raw deltas, clicks, and scroll events with hardware timing. The ground truth of how a human actually moves through a world.

    Camera telemetry

    3D camera position, rotation, field of view, and velocity extracted directly from the game engine.

    World state physics

    Engine-level physics data: object positions, collisions, velocities, and environmental state. Structured as synchronized JSON.

    Game state + action events

    Game-state changes and in-engine action events share one synchronized stream. Action-labeled timelines tag what is happening, where, and when across 1,000+ activity types.

    More to come

    The capture SDK keeps growing. New modalities ship as engines expose them - ask us what is on the roadmap.

    Real world models cannot be built off hyper-sensationalized 10-second clips of headshots uploaded online. You need the diversity. The boring. You need everything - not the 1% that makes it online, or even the 10% that video-game players do. You need the 90% of things they would never do, but CAN do.

    This is why Origin Lab's data is different.

    A top frontier model researcher

    03/How it's captured
    Curated,
    not scraped.

    Signals are only as good as the hours behind them. Our data is captured by professional players following structured task lists, planned for diversity per hour, and monitored in real time. Every recording earns its place in the corpus.

    1. 01

      Title-specific task lists

      Granular, per-game instructions. Complete every crafting loop. Drive backwards through every biome. Play only at night. Interact with every NPC type.

    2. 02

      Diversity over volume

      Each run is planned to cover the widest range of combinatorial actions, environments, and edge cases per hour. Models see the full breadth of what a game world can produce.

    3. 03

      Real-time monitoring

      Webcam-based engagement tracking and input-activity detection. AI-driven coverage planning steers future runs toward gaps. No idle time enters the corpus.

    4. 04

      Custom capture by request

      Share your spec or RFP. We design capture plans around your requirements and execute across any title in our catalog.

    04/The platform
    Studios partner.
    Origin Lab enriches.
    Labs train.

    Signals and sessions only matter if they reach you cleanly. The platform threads rights holders, our capture team, and your training pipeline through one license, one schema, one API.

    Who partners with usRights-cleared content from licensed sources
    AAA game studiosIndie publishers3D studiosVFX housesMocap studios
    LIVEOrigin Lab pipeline

    Capture, enrich, align -frame by frame.

    01
    Verify rights
    Legal
    02
    Read the scene
    Schema
    03
    Record gameplay
    Recorder
    04
    Enrich with AI
    Labels
    05
    Stream to your stack
    API
    Who we powerTraining data for frontier model teams
    AI foundation labsWorld-model buildersVideo generationGame-AI companiesRobotics & embodied AISim-to-real researchPlanning & eval teams

    05/Library + spec
    Explore what we have.
    Capture what you need.

    Twenty-plus publisher partners. Fifty-plus licensed titles across a wide range of game genres, plus 3D environments. Growing daily. Don't see your target genre, mechanic, environment, or modality? We capture to spec.

    20+Publisher partners
    50+Licensed titles
    100%Captured for AI training specs

    Games

    Engine-level capture from AAA and indie titles. First-party exports with action labels, camera telemetry, depth, normals, and game state.

    By size
    AAAAAIndie
    By genre
    FPSOpen-worldSurvivalRPGActionRacingCo-opSimulation

    3D Environments

    Worldbuilder libraries with full scene graphs. USD, FBX, glTF assets ready for simulation, training, and synthetic data.

    By type
    UrbanNaturalIndustrialInteriorFantasySci-fi

    Need something we don't have yet?

    Share a spec, an RFP, or just your training goal. We design capture plans around your requirements - across any title in our catalog, or one we license for you.

    Request a custom capture

    06/Engine vs. web
    Engine data direct from the source, not scraped or inferred.

    This is what scraping the web can't produce. Web pixels are volume and legal exposure. Engine capture is signal density, time-aligned modalities, and a clean rights trail.

    Dimension
    Origin Lab
    Scraped video
    Source
    Licensed engine data
    Web pixels
    Modalities
    10+ aligned
    RGB only
    Metadata
    20+ categories
    None
    Rights status
    100% cleared
    Unclear / risky
    Turnaround
    24h sample
    Instant, unusable
    Audit trail
    Per-frame provenance
    None
    "Scraping the web for video gives you pixels without physics. Origin Lab gives you the underlying world."
    - Research lead, frontier AI lab

    Build on real worlds.

    Sample pack in 24 hours. Same-day API access for evals. Whether you're training a frontier model or wiring up a single eval, start with one signed agreement and one cURL.

    See the API