What is human-in-the-loop data collection?
Demonstrations recorded from scratch teach a policy how an expert behaves in the states an expert visits. Once the policy runs on its own, it drifts into states the demonstrations never covered, and small errors compound. Human-in-the-loop collection closes that gap by letting the policy act, watching for mistakes, and having the operator intervene, so the recorded corrections cover the states the policy reaches on its own.
Variants differ in who decides when to intervene. In DAgger, an expert labels the states the policy visits with the correct actions. In human-gated versions such as HG-DAgger, the operator watches and takes control only when needed, and the dataset marks which steps were human interventions. LeRobot supports this pattern through its rollout tooling, which can record autonomous steps and human corrections in the same dataset format used for training.
Key takeaways
- A person supervises a running policy and takes over when it fails, and those corrections become training data.
- New data concentrates on the states the policy reaches on its own, which demonstrations alone miss.
- Datasets record which steps were interventions, so teams can weight, filter, or analyze them.
How it works
The policy runs on the robot while an operator monitors it, usually with a leader arm or another teleoperation device ready. When the operator takes over, control switches from the policy to the human, and the recorded steps are marked as interventions. After the session, the corrected episodes are added to the training set and the policy is retrained, often over several rounds.
Why it matters
Human-in-the-loop data collection matters because it spends expensive human time where it does the most good. Instead of recording more demonstrations of situations the policy already handles, operators supply corrections for its failure cases. It also connects evaluation to data collection: every rollout is both a test of the current policy and a source of new training data.
Frequently asked questions
What is DAgger?
DAgger, short for dataset aggregation, is an imitation learning algorithm that repeatedly runs the current policy, asks an expert to label the states it visits with the correct actions, and retrains on the combined data. Human-in-the-loop data collection applies the same idea with a person correcting the robot in real time.
How are interventions stored in a dataset?
Datasets typically include a per-step flag marking whether the human or the policy was in control. Teams use it to train on corrections, measure how often intervention was needed, or analyze where the policy fails.
Related terms