Datachain

Your coding agent can now work with massive video and sensor data

Claude Code, Codex, or Cursor: search video/sensor data, curate datasets, run code on millions of files across hundreds of machines — data stays put.

Try DataChain

Trusted by startups and Fortune 500 companies

  • DeGould
  • Kibsi
  • Hugging Face
  • Mantis
  • Motorway
  • Inlab
  • Pieces
  • Sicara
  • UBS
  • SumerSports
  • UK Hydrographic Office
  • XP Inc.
  • Cyclica
  • Aicon
  • Billie
  • Papercup

One agent, from raw recordings to the dataloader

Our pedestrian detector regressed in night rain. Find the failures, figure out what they have in common, and build a dataset we can evaluate the fix against.
One agent. Data flows from Collection (the fleet) through three stages handled by DataChain: Ingestion, compute what nobody labelled; Curation, assemble and version the set; and Loading, stream it into the training set. One store, many questions. Collection and Training (your model) are outside DataChain.

One question. Four
different kinds of work

  • Finding the failures reads existing recordings to locate model errors.

  • Finding common patterns extracts new signals shared by the failures.

  • Building the eval set saves selected examples as a new dataset.

  • Evaluating the fix hands the dataset to training to test improvements.

Real data. Real research tasks

Failure investigation

From a regression to the exact recordings and failure cases behind it, then uncover the shared signals that explain what went wrong.

Dataset curation

From scattered files, recordings, labels, and metadata to a reproducible, versioned dataset ready for evaluation or model training.

Large-scale processing

From a local Python prototype to the same workload running efficiently across millions of files and hundreds of machines.

Follow-up research

From an expensive first analysis to new questions that reuse existing results, code, datasets, and state instead of starting from scratch.

What our customers say

  • DataChain added real value to our workflows - versioned datasets, automated ETL, and MLOps, all in Python. If you need a data management layer on top of cloud storage, give it a try.

    Nikhilesh Saggere

    Lead Engineer, Alps Alpine Europe

  • What surprised me was how easily researchers adopted DataChain - data tools are usually hard for non-engineers. What surprised me more was when hardware and QA started asking for access too.

    Sharon Kohen

    Principal Data Engineering, brain.space MobiClocks

  • We realized we were solving a problem we shouldn't be solving. With DataChain, what used to require data engineers is now handled seamlessly by researchers - and the whole team moved to the next level.

    Yoni Svechinsky

    Director of Research, brain.space

Remember what worked.
Run what comes next

Data Memory: Datasets, code, lineage, and dependencies. Execution: Run Python over millions of files on hundreds of machines.

Claude Code, Cursor, and Codex connect to the Physical AI Data Runtime. Data Memory retains datasets, code, lineage, and dependencies. Execution runs Python at scale, with checkpoints and incremental processing. The runtime connects to your object storage: S3, GCS, or Azure.

Every run creates new state the next task can reuse.

Your data never moves. Everything else does

The compute runs next to your buckets. The state data warehouse lives there too. Nothing is copied out, and nothing is read across the internet.

Your Claude Code connects through MCP to the DataChain control plane, which manages keys, limits, audit, and usage stats. Your cloud account holds your buckets of raw recordings, your VPC for CPU and GPU enrichment passes, and the state: datasets, schemas, code, and lineage.

SOC 2 Type II compliant. We hold the keys, the limits, the audit and usage stats. Everything else is yours, including the state your team builds.

Give your coding
agent a real data task

Run it from Claude Code, Codex, or Cursor - on the data already in your cloud.

Try DataChain