FAQ · Common questions

    Questions, answered.

    Everything you need to know about licensing world data, working with Origin Lab, and how we protect both buyers and rights holders.

    For AI labs
    What do you deliver?
    Best-in-class multimodal training data, built from the ground up for frontier AI. We have deep expertise in video game worlds and 3D environments, all with AI-enriched metadata, structured packaging, and richer signals like input telemetry and engine data where available.
    Can you fulfill built-to-spec requests?
    Yes. Share the spec and target distributions. We return a sample quickly, iterate once, then deliver at scale.
    How do you handle rights and compliance?
    Every source is 100% licensed with verified ownership and defined usage scope before any capture begins. We maintain a full chain of custody from source through capture, enrichment, and delivery. Rights holders participate in the value their content helps create, with clear terms, IP protections, and structured compensation. Audit-ready artifacts (license agreements, dataset cards, and provenance metadata) ship with every export.
    Can we use this for training, fine-tuning, and evals?
    Usage depends on the license. We structure terms explicitly so your intended use is covered without ambiguity.
    How is Origin Lab different from Scale AI, Appen, or Bright Data?
    Most data providers start with the open web and work backward. Scraping or aggregating content first, then trying to clean, label, and license it after the fact. Origin Lab works in the opposite direction. Every source is licensed before capture begins. Every dataset includes full chain of custody, engine-level telemetry, and AI-enriched metadata that goes far deeper than tags or labels. We specialize in interactive worlds: video games and 3D environments. Not generic web content.
    What does it cost?
    Every dataset is priced using our proprietary valuation engine, which estimates training utility for your specific use case. Pricing is tied to volume and utility, not flat rates. Reach out with your needs and we’ll scope a proposal quickly. Most buyers have a first conversation and sample within days, not weeks.
    Can I see sample data before committing?
    Yes. Reach out with your use case and we’ll share sample schemas, metadata examples, and representative clips so you can evaluate fit before any commitment. We want you to see exactly what you’re getting: the structure, the depth, and the signals. Before we scope anything further.
    How do we start?
    Request access with your needs and timeline. We will share a sample schema, a dataset card, and availability aligned to your spec.
    What are the capture data specs?
    Video is captured at 1080p and up to 4K (3840×2160) at 30–60fps. Audio is recorded at 48kHz stereo. Webcam footage, when included, is captured at 1080p. All captures include synchronized timestamps across all signal streams.
    What are the minimum and maximum clip lengths?
    Clip lengths typically range from 10 seconds to 30 minutes depending on the content type and use case. For gameplay capture, sessions can be up to 2 hours before segmentation. We can adjust min/max lengths to match your training pipeline requirements.
    What metadata format do you use?
    Metadata is delivered in JSON and Parquet formats. Every clip includes a structured metadata record with signals, tags, scene annotations, input telemetry (keyboard, mouse, camera), engine telemetry (position, velocity, physics), and frame-level attributes (resolution, weather, biome, time of day). A dataset card accompanies every delivery.
    What signals are included beyond video?
    Ten modalities today, with more coming: pre-HUD RGB, post-HUD RGB, depth maps, surface normals, audio, keyboard inputs, mouse inputs, camera telemetry, world state physics, and action-event extraction. Frame metadata (weather, lighting, biome) and scene-level annotations ride alongside. Not all modalities are available for all content. Availability depends on engine access.
    What’s included in each catalog type?
    Video Games: high-fidelity gameplay capture with input telemetry, frame-level metadata, and action annotations (MP4, custom encoding). 3D Worlds: licensed environments, scenes, and assets with scene-level annotation exports (glTF, FBX, custom).
    What export formats do you support?
    MP4 for video, Parquet and JSON for metadata, glTF/USD/FBX for 3D assets. We can package custom formats for specific training pipelines upon request.
    How is data versioned?
    Every delivery includes a schema version, dataset ID, and dataset card. When datasets are updated (new captures, corrections, or enrichment), we issue versioned updates so your pipeline can track changes cleanly.
    For IP holders
    What kinds of catalogs are a fit for Origin Lab?
    We work with AAA video game studios, indie developers, 3D artists, and content distributors. If you have rights-cleared content across video games or 3D environments with clear ownership, we’d like to hear from you.
    Do I keep ownership of my IP?
    Yes. We license usage rights under defined terms. You retain ownership of the underlying IP, and licensing scope is explicit and auditable.
    How do you verify rights?
    We require licensor authorization and rights verification before any capture begins. Every license is scoped, and IP protections are built in.
    How do payouts work?
    We structure deals to fit the catalog and buyer demand, with negotiated licensing fees and IP protections built in. Terms vary by modality, exclusivity, and scope.
    What is the process to get onboarded?
    We start with a free evaluation of your content, at no cost and no commitment. From there the process is: evaluate your catalog, capture and enrich the content, list it on our platform, connect it with buyers, and you get paid. You retain full ownership of your IP, with clear terms and protections at every step.
    How do you evaluate my content and estimate hours?
    We use a proprietary valuation engine that scores your content across world size, environment diversity, interaction density, physics fidelity, and replay potential. The model estimates how many unique, high-value training hours your content can produce based on its structure, richness, and downstream utility for AI models.
    How do you price those hours?
    Pricing is driven by the same valuation engine. Each title is scored for training utility across multiple AI use cases, and pricing scales with volume. Higher-value content commands higher per-hour rates, and prices adjust as cumulative hours increase. The result is transparent, data-driven pricing tied directly to the quality and utility of your content.
    For researchers
    Do you work with research teams?
    Yes. We’re actively collaborating with AI researchers from Oxford, Google Research, and others, providing training data, access, and support to drive breakthroughs in Artificial World Intelligence™. We welcome new collaborations where rights, publishing constraints, and intended use are clear.
    What makes Origin Lab useful for world and spatial research?
    We can pair multimodal content with richer structure such as 3D environments and gameplay capture signals, enabling experiments beyond raw video alone.
    Can I publish examples from the datasets?
    Publishing constraints depend on the rights. We can often support research outputs with approved examples or derived artifacts rather than raw content.
    Are you building synthetic and generative systems too?
    Yes. Origin Research is actively working across five tracks: generative world-building pipelines, synthetic human players for capture and QA, scene intelligence for indexing and retrieval, world value scoring for training utility, and automated data quality and AI QA. The goal is to improve the fidelity, efficiency, and controllability of future datasets.
    How do I engage?
    Reach out with your research objective, modality needs, and what you plan to publish. We will suggest the best path and constraints.
    Legal & security
    How do you verify licensed sources?
    We require licensor authorization and rights verification before any capture begins. Every license is scoped, and IP protections are built in. Audit-ready artifacts — license agreements, dataset cards, and provenance metadata — ship with every export.
    Do you offer indemnification?
    Yes. We offer indemnification as part of our enterprise licensing terms. Every dataset ships with full provenance documentation and a clear chain of custody to back it up.
    Where is data stored?
    Data is stored in SOC 2-compliant cloud infrastructure with encryption at rest and in transit. We support regional storage requirements and can work with your security team to meet specific compliance needs.

    Still curious?

    If you didn’t find what you were looking for, reach out directly. We respond to every message.