Game Recordings v3: synchronized multimodal gameplay
Human gameplay captured in-engine at 1080p and 60 FPS, with RGB, depth, normals, audio, actions, camera, and telemetry aligned on one shared frame clock.

Twenty-plus publisher partners. Fifty-plus licensed titles. Every recording carries 10 modalities frame-locked on one clock - RGB before and after the HUD, depth, normals, audio, inputs, camera, physics, and action labels - delivered through a single API, with more on the way.
Explore the data, methods, and checkpoints behind our latest research. The complete release record lives on Hugging Face.
Origin Lab on Hugging FaceSwipe to explore all four releases
Human gameplay captured in-engine at 1080p and 60 FPS, with RGB, depth, normals, audio, actions, camera, and telemetry aligned on one shared frame clock.

A depth model trained only on video game frames beats purpose-built synthetic data on real KITTI photographs, with 4x less training data.

Eight PCs. Eight players. One frame clock. Every prior multi-view gameplay dataset was rendered from a replay; this one was captured live on independent machines.

Released depth checkpoints make the Game-Depth result directly inspectable and reproducible.
Every release ships under the Origin Lab Data License with a 90-day internal evaluation track for labs. Commercial licensing requires a direct agreement. How release terms apply
Real licensed gameplay, captured in-engine at a 60 fps CFR standard. Events, Camera, Input, and Feed are different views into one synchronized record: ten modalities, game state, and in-engine action events on the same timeline.
What you just watched is 10 modalities captured at the engine level - depth, normals, telemetry, inputs, and action labels locked to the same clock as the rendered frame, at the source. They cannot be inferred from video alone.
Pristine gameplay at up to 4K/60fps. HUD, menus, and UI stripped at the engine level before recording - your model trains on the raw 3D world, not on screen capture.
The same frame exactly as the player saw it, HUD and interface included. For agents and models that need to read the screen, not just the world.
Per-pixel depth straight from the renderer, not estimated after the fact. The same maps behind the slider in the demo above.
Per-pixel surface orientation from the engine. Geometry, lighting, and material structure for reconstruction and 3D understanding.
In-game audio captured as separate tracks. Dialogue, environmental sound, and effects at full fidelity.
Exact key sequences straight from hardware, with press and release timing. Not inferred from video.
Raw deltas, clicks, and scroll events with hardware timing. The ground truth of how a human actually moves through a world.
3D camera position, rotation, field of view, and velocity extracted directly from the game engine.
Engine-level physics data: object positions, collisions, velocities, and environmental state. Structured as synchronized JSON.
Game-state changes and in-engine action events share one synchronized stream. Action-labeled timelines tag what is happening, where, and when across 1,000+ activity types.
The capture SDK keeps growing. New modalities ship as engines expose them - ask us what is on the roadmap.
Real world models cannot be built off hyper-sensationalized 10-second clips of headshots uploaded online. You need the diversity. The boring. You need everything - not the 1% that makes it online, or even the 10% that video-game players do. You need the 90% of things they would never do, but CAN do.
This is why Origin Lab's data is different.
Signals are only as good as the hours behind them. Our data is captured by professional players following structured task lists, planned for diversity per hour, and monitored in real time. Every recording earns its place in the corpus.
Granular, per-game instructions. Complete every crafting loop. Drive backwards through every biome. Play only at night. Interact with every NPC type.
Each run is planned to cover the widest range of combinatorial actions, environments, and edge cases per hour. Models see the full breadth of what a game world can produce.
Webcam-based engagement tracking and input-activity detection. AI-driven coverage planning steers future runs toward gaps. No idle time enters the corpus.
Share your spec or RFP. We design capture plans around your requirements and execute across any title in our catalog.
Signals and sessions only matter if they reach you cleanly. The platform threads rights holders, our capture team, and your training pipeline through one license, one schema, one API.
Twenty-plus publisher partners. Fifty-plus licensed titles across a wide range of game genres, plus 3D environments. Growing daily. Don't see your target genre, mechanic, environment, or modality? We capture to spec.
Engine-level capture from AAA and indie titles. First-party exports with action labels, camera telemetry, depth, normals, and game state.
Worldbuilder libraries with full scene graphs. USD, FBX, glTF assets ready for simulation, training, and synthetic data.
Share a spec, an RFP, or just your training goal. We design capture plans around your requirements - across any title in our catalog, or one we license for you.
This is what scraping the web can't produce. Web pixels are volume and legal exposure. Engine capture is signal density, time-aligned modalities, and a clean rights trail.
Sample pack in 24 hours. Same-day API access for evals. Whether you're training a frontier model or wiring up a single eval, start with one signed agreement and one cURL.