Human-in-the-loop data collection

Human-in-the-loop data collection is a way of gathering robot training data in which a person supervises a learned policy as it runs and takes over when it starts to fail. The corrections are recorded and added to the training data, so new data concentrates on the situations where the policy struggles. The approach builds on the DAgger family of imitation learning methods.

What is human-in-the-loop data collection?

Demonstrations recorded from scratch teach a policy how an expert behaves in the states an expert visits. Once the policy runs on its own, it drifts into states the demonstrations never covered, and small errors compound. Human-in-the-loop collection closes that gap by letting the policy act, watching for mistakes, and having the operator intervene, so the recorded corrections cover the states the policy reaches on its own.
Variants differ in who decides when to intervene. In DAgger, an expert labels the states the policy visits with the correct actions. In human-gated versions such as HG-DAgger, the operator watches and takes control only when needed, and the dataset marks which steps were human interventions. LeRobot supports this pattern through its rollout tooling, which can record autonomous steps and human corrections in the same dataset format used for training.

Key takeaways

  • A person supervises a running policy and takes over when it fails, and those corrections become training data.
  • New data concentrates on the states the policy reaches on its own, which demonstrations alone miss.
  • Datasets record which steps were interventions, so teams can weight, filter, or analyze them.

How it works

The policy runs on the robot while an operator monitors it, usually with a leader arm or another teleoperation device ready. When the operator takes over, control switches from the policy to the human, and the recorded steps are marked as interventions. After the session, the corrected episodes are added to the training set and the policy is retrained, often over several rounds.

Why it matters

Human-in-the-loop data collection matters because it spends expensive human time where it does the most good. Instead of recording more demonstrations of situations the policy already handles, operators supply corrections for its failure cases. It also connects evaluation to data collection: every rollout is both a test of the current policy and a source of new training data.

Frequently asked questions

What is DAgger?

DAgger, short for dataset aggregation, is an imitation learning algorithm that repeatedly runs the current policy, asks an expert to label the states it visits with the correct actions, and retrains on the combined data. Human-in-the-loop data collection applies the same idea with a person correcting the robot in real time.

How are interventions stored in a dataset?

Datasets typically include a per-step flag marking whether the human or the policy was in control. Teams use it to train on corrections, measure how often intervention was needed, or analyze where the policy fails.

Related terms

black and white photo of Jesse Mostipak
Jesse Mostipak
Director of Growth
Jesse Mostipak is the Director of Growth at Voxel51, where the work is helping humans find and trust what the brand knows, and teaching the Google knowledge graph and the LLMs answering on their behalf to do the same. That question, how knowledge gets built inside a system, is one Jesse has been chasing for years. Earlier versions of it ran through a New York City high school science classroom, data science and machine learning, and developer relations at Kaggle, Posit (formerly RStudio), and Baseten. The answer doesn't change much depending on whether the learner is a teenager, a software engineer, or a knowledge graph. Jesse holds a Master's in Education from CUNY Hunter College. LinkedIn
See all articles by Jesse Mostipak
Last updated October 8, 2026

Building visual or physical AI?

Let's talk.