You Can’t Watch 10,000 Episodes. Measure These Instead: Curation Metrics for LeRobot Datasets
Oct 8, 2026
•
27 min read
Author
Harpreet Sahota
Hacker-in-Residence
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.
A field guide to the metrics worth computing on a LeRobot dataset before you train a policy or fine-tune a vision-language-action (VLA) model: what each one measures, when to trust it, where it misleads you, and how to run the whole workflow in FiftyOne.
You can’t watch all of it. At 30 seconds per episode, 10,000 episodes is about 83 hours of video, and that’s before you open a second camera. So you scrub through a handful, they look fine, and you hit train.
Then the policy hesitates at the first grasp. Somewhere in there is an episode with a frozen wrist camera. Another has 11 seconds of dead time before the arm moves. A third has the task string "Hold". A fourth has a video track that runs 0.4 seconds behind the actions. None of these show up when you watch one episode at a time because none of them is visible in a single episode. They show up when you compare an episode to its peers.
That’s the gap this post is about. A viewer shows you data. Curation tells you which episode to open next.
This guide walks through the metrics that do that job for LeRobot and VLA datasets. For each one, it covers what it measures, why you’d compute it, when to use it, how to read it, and where it lies to you. It also shows where each metric lives in a FiftyOne panel, because a metric you can’t filter, tag, and export is just a number in a notebook.
The post isn’t about any one dataset. The metrics apply to anything in LeRobot v3 format.
Key takeaways
Curation metrics answer six different questions: is the episode broken (integrity), is the motion clean (smoothness), is it wasting frames (time), did the robot do what was commanded (tracking), can you trust the cameras (vision), and can a model learn from the instruction (language). Start with the cheapest.
Most of these metrics need no model and no GPU. Integrity checks read metadata and Parquet. Motion, time, tracking, gripper, and consistency metrics are signal math on action and observation.state. Only the camera metrics decode video, at roughly 0.4 s per camera per episode.
Compare episodes to their own task group, not to the whole dataset. A fold-a-shirt episode and a pick-a-cube episode have different “normal” motion, and a pooled score will rank one task as the problem.
Correlated metrics should share one vote. Five time metrics that all fire on the same pause are one problem, not five.
The published evidence is mixed. Spectral arc length (SPARC) smoothness filtering raised RoboMimic success by 16% using one-sixth of the data (RINSE, arXiv:2604.23000). A separate audit (arXiv:2606.10229) found 5 of 7 curation metrics exploited episode length, and that detection accuracy and policy success were uncorrelated. Use scores to order a review queue, and validate with a training run before you delete anything.
FiftyOne’s role is the workflow around the metrics: one sample per episode, a panel that re-ranks as you filter, click-to-filter charts, temporal tags on the episode timeline, and a lossless export of the kept subset as a LeRobot v3 dataset.
First, the LeRobot v3 format, in one paragraph
A LeRobot v3 dataset doesn’t store one file per episode. Frames from many episodes are concatenated into shared Parquet files (data/chunk-000/file-000.parquet), and each camera’s video is concatenated into shared MP4s. Metadata in meta/info.json holds the schema and fps, meta/stats.json holds normalization statistics, and meta/episodes/ holds one row per episode with its row range and, per camera, the video time window.
Two consequences matter for curation. First, “episode N” points into shared files, so bookkeeping can drift without anything crashing. Second, the timestamp column is written from the nominal fps, not from a logged clock, so classic sensor-health checks such as jitter and clock drift don’t apply. You check structure and alignment instead.
FiftyOne’s native fo.types.LeRobotDataset (FiftyOne 1.22 and later) imports each episode as one multimodal sample, with a synchronized episode viewer on top (API reference). It reads only the metadata up front and pulls Parquet rows and video frames on demand, which is what makes dataset-level scoring practical.
What you have to tell the tool
Nothing about your robot should be guessed. A 14-dimensional action vector could be two 7-degree-of-freedom (DoF) arms, or one arm plus a hand, or end-effector deltas. If a tool silently picks, every smoothness score downstream is wrong in a way you can’t see.
So the compute form asks for five things:
Which arrays are state and action. They default to observation.state and action when the dataset uses the LeRobot standard names, and stay unset otherwise.
What the action means. Is it absolute joint positions in the same space as the state, or something else (deltas, end-effector poses)? This is a yes/no question, and it only appears when the two arrays have the same size.
Robot and joints. Single or dual arm, whether there’s a multi-joint hand, and which joints are the arm and which are the gripper. You pick them by name, with each joint’s observed range beside it.
Gripper open direction. Does a high value mean open or closed?
Cameras. Which video cameras to score. None are selected by default, because decoding is the slow part.
Anything you leave unanswered switches off the metrics that need it, and the form lists exactly which. It never blocks the run.
Good to know. Why does “what the action means” matter so much? Tracking, acceleration-spike, and joint-limit metrics compare the commanded action to the achieved state. That comparison only means something if both live in the same joint space. If your action is end-effector deltas, a gap between the two is just a unit mismatch, so those metrics stay off.
The LeRobot curation metrics
Every metric below follows the same pattern: what it measures, why it matters, when to use it, how to read it, and the pitfall to watch for. The harder ones get an intuition-level explanation.
One convention first. Inside a score, every metric is oriented so that higher means worse. The tables and charts show each metric’s raw direction, and the scoring section explains how they combine.
1. Integrity: is this episode broken?
Integrity checks don’t say an episode is worse. They say it’s broken. They produce pass, warn, or fail flags and counts. They never enter the ranking score, and they get their own verdict. They’re also the cheapest checks you’ll run, since they read metadata and Parquet and decode nothing.
Run them first. There’s no point ranking the smoothness of an episode whose rows don’t match its video.
Figure 1. A v3 episode is a row range in a shared Parquet file plus a time window in a shared MP4 per camera. Each integrity check catches a case where a pointer stopped matching the data. Illustrative values.
Check
What it catches
frame_gaps
Non-contiguous or duplicated frame_index values. Any nonzero count means a broken episode.
length_mismatch
The episode’s declared length disagrees with the actual number of Parquet rows. Metadata and data disagree.
video_window_mismatch
Per camera, the video window duration disagrees with length/fps. Video and actions differ in length.
nonfinite_values
NaN or inf in action or state. Any nonzero count breaks training.
too_short
Episodes under a minimum duration or frame count. Catches one-frame and near-empty episodes.
schema_mismatch
When you merge sources: differences in feature shapes and names, fps, units, camera keys, or resolution. Flags a dataset that won’t collate cleanly.
The six integrity checks and the broken-episode case each one catches.
timestamp_dev (the largest deviation of consecutive timestamps from 1/fps) is reported as information only. Because LeRobot writes timestamps from fps, it’s usually about zero and doesn’t affect the verdict.
Why it matters. The Hugging Face team’s write-up on community datasets lists these defects as recurring: episodes with one or very few frames, Parquet files deleted without reindexing, and inconsistent action or state dimensions for the same robot. Each one fails silently until a dataloader hits it, or worse, doesn’t fail at all.
Pitfall.video_window_mismatch compares durations, so it can’t see a constant offset. An episode whose video starts 400 ms late passes. That’s what the camera lag metric later in this guide is for.
2. Motion: is the motion clean?
Motion metrics measure the shape of the movement an operator commanded. This is where most of the curation literature lives, and it’s also where the most care is needed, so the preprocessing matters as much as the metric.
What signal it reads. The action array over the joints you put in the Arm (and Hand) fields. action is what a policy learns to output. On leader-follower rigs such as SO-100, Koch, and ALOHA, it’s the operator’s hand, not the robot’s arm.
How the signal is prepared.
Each joint is range-normalized before speeds are combined, so a joint with a large range doesn’t dominate and units don’t matter.
Gripper and hand dimensions are excluded from arm speed. A 22-joint dexterous hand is scored as its own signal so it never swamps a 7-joint arm.
The low-pass cutoff is min(10 Hz, 0.4 × fps). At 30 fps that’s 10 Hz, well away from the Nyquist frequency of 15 Hz.
Metrics run on fixed-length windows because some are sensitive to episode duration. Each scored metric also keeps a worst-window value, so one bad stretch isn’t averaged into a fine episode.
sparc: how many corrections does the motion contain?
Measures. The spectral arc length of the arm’s speed profile.
Intuition. Take the arm's speed over time. A clean reach is one smooth bump: speed rises, peaks, falls. Take that curve’s frequency spectrum and trace along it. A single bump has a compact spectrum, so the traced path is short. Every extra correction, hesitation, or stop-and-go adds frequency content, and the path gets longer. SPARC is the negative of that path length, so smooth motion sits near 0, and fragmented motion goes very negative.
In its standard form, SPARC is the negative arc length of the normalized speed spectrum up to a cutoff frequency:
where V̂(ω) is the magnitude spectrum of the speed profile, normalized by its peak. You don’t need to compute that by hand to use the metric. The picture matters: more corrections mean a longer trace.
Why it matters. It’s the most direct validation in the literature. RINSE (arXiv:2604.23000) reports that SPARC-based filtering gave 16% higher RoboMimic success using one-sixth of the data. It also reports SPARC as about 10 times more noise-robust than log dimensionless jerk (LDLJ).
Use it when. You want one primary smoothness number per episode. It’s scale-invariant, so it works across arms of different sizes.
Read it. Closer to 0 is smoother. Very negative means fragmented, hesitant motion.
Pitfall. SPARC responds to stop-and-go motion, not to additive white noise. A jittery signal from a noisy encoder can still look smooth to SPARC. That’s what the next metric is for.
ldlj: how much jerk, once you account for size and speed?
Measures. Log dimensionless jerk of the speed profile. Jerk is the rate of change of acceleration. It’s integrated over the movement, made dimensionless using the movement’s duration and peak speed, then log-transformed.
Why it matters. It catches what SPARC doesn’t: high-frequency noise and abrupt changes.
Use it when. You suspect noisy teleop hardware or a bad leader-arm encoder, alongside sparc.
Read it. Closer to 0 is smoother.
Pitfall. It’s noisier than SPARC and duration-sensitive, which is why the windows are fixed-length and why it carries half weight in the score.
Figure 2. The same two reaches, moved three ways. SPARC: smooth -2.57, stop-and-go -5.25, jittery -2.56. LDLJ: smooth -8.87, stop-and-go -14.45, jittery -14.71. SPARC separates stop-and-go from the other two. LDLJ penalizes both stop-and-go and noise. Synthetic signals at 30 fps. Values computed from them.
sparc_phase: smoothness within each phase of the grasp (opt-in)
Measures. SPARC computed inside each gripper-delimited phase (approach, grasp, transport, and release), rolled up by median.
Why it matters. Contact is inherently jerky. When the gripper closes on an object, speed changes abruptly for good reasons. Scoring the whole episode as one signal counts those transitions as roughness. Scoring per phase stops that.
Use it when. Your task has clear grasp and release events, and you’ve declared gripper joints and the open direction.
Read it. Same as sparc.
Pitfall. It adapts the gripper-phase idea from RINSE and isn’t the paper’s exact method. It’s covered by synthetic tests only. It skips phases with too few samples.
Two more motion metrics exist for review only and are never scored. jerk_rms is the root mean square (RMS) jerk after a low-pass filter. It’s noise-sensitive and correlated with the other metrics. psd_lf_hf is the log ratio of low- to high-frequency power in the speed profile, and it’s unreliable at LeRobot frame rates, where tremor sits near Nyquist.
3. Time efficiency: is the episode wasting frames?
Time metrics measure dead time. They read the commanded motion (action), with an idle threshold set as a multiple of the episode’s own moving speed. That per-episode threshold means a slow, careful operator isn’t flagged just for being slow.
Metric
Measures
Read it
idle_lead_s
Seconds of stillness at the start
High means dead time before the task begins. Also stores a suggested trim start (keep_from_s).
idle_trail_s
Seconds of stillness at the end
High means a long dead tail. Stores a suggested trim end (keep_to_s).
longest_pause_s
Longest idle run that isn’t at either end
High means a mid-task stall or hesitation.
end_motion_ratio
Speed in the last second relative to typical moving speed
High means the episode probably ended mid-motion.
length_z
Robust z of episode duration within its task group
Scored as abs(z). Positive means unusually long (struggle). Negative means unusually short (abandoned).
Time efficiency metrics, what each one measures, and how to read a high value.
Figure 3. Left: one episode's arm speed with the three idle stretches the time metrics measure, and the keep window they suggest. Right: length_z places an episode's duration against its task group. Synthetic signals.
Why it matters. Two published pipelines treat this as signal. OpenVLA removed all-zero no-op actions from its training data. SCIZOR’s removed samples included “manipulation failures, slow motion, and pauses.” Both pipelines treated dead frames as removable, and either way they cost training compute.
Use it when. Your data comes from human teleop, where operators pause to think, reposition, or wait for a timer.
Pitfalls.
Trailing idle and end_motion_ratio get half weight, because fixed-duration recording setups produce a dead tail from the timer itself.
A pause next to a regrasp counts as recovery and is excluded from longest_pause_s. A hesitation that precedes a retry is the operator fixing a problem, not wasting time.
idle_frac (the overall fraction of idle frames) is shown for context and never scored. A high value can be a legitimate pause. path_length_z (opt-in) is also never scored, since it correlates with length_z.
length_z, path_length_z, and idle_frac depend on episode length by construction. Only length_z is scored, and it shares one vote with the idle metrics. That design decision comes straight from the length-confound audit discussed later.
Trimming is a suggestion.keep_from_s and keep_to_s store a proposed range. Nothing is written back to the dataset.
4. Tracking and contact: did the robot do what was commanded?
On a leader-follower rig, action is where the operator’s leader arm is, and observation.state is where the follower robot actually is. When the two agree, the robot is doing what it was told. When they diverge, something physical happened: a collision, a stall, a slipping joint, or a following arm that couldn’t keep up.
These metrics read the state, so they run only when you’ve declared the action to be absolute joint positions in the state space.
Figure 4. The follower trails the leader by a few frames (right), then stalls while the command keeps moving. The residual after lag alignment lights up exactly where the stall is, and the catch-up produces an acceleration spike. Synthetic signals.
track_resid: how far is the follower from the command?
Measures. The per-joint, normalized gap between action and state after lag alignment, with each joint’s median offset removed first. Arm joints only.
Why it matters. Calibration offsets are constant and harmless. You want the part of the gap that varies, where contact and slippage live.
Read it. High means the robot isn’t reaching the commanded position. It also raises interval flags that become spans on the timeline.
track_lag_ms: how late is the follower?
Measures. The delay between command and response, estimated by sliding one signal against the other and picking the offset with the highest cross-correlation:
lag = argmax_k corr( action[t], state[t + k] )
Why it matters. Mostly it’s a hardware property. A servo with a slow control loop lags every episode by about the same amount.
Read it. It’s summarized per dataset or session, and only episodes far from the dataset median are flagged. It never enters a score.
accel_spike_frac: sudden jolts
Measures. The fraction of frames where any joint’s state acceleration exceeds median + k × MAD (median absolute deviation), with an absolute floor.
Why it matters. It’s a proxy for collisions and contact. The absolute floor stops idle joints from turning every frame into a spike.
joint_limit_frac (opt-in): operating near the edge
Measures. The share of (frame, joint) pairs within a margin (default 2% of range) of a joint limit.
Pitfall. It uses your limits if you give them, and otherwise falls back to meta/stats.json min and max. That’s a weak proxy, because observed extremes always sit at the edge of the observed range.
Keep it honest. These metrics have no published downstream validation. Treat them as sanity checks. Related work (RoboDrop, arXiv:2609.10021) shows observation-action temporal misalignment is a major corruption that action-only methods handle poorly, and a teleoperation paper (arXiv:2605.26349) flags stalls and operation near joint limits, but its user study had 3 operators. That’s weak evidence, not a validated benchmark.
5. Gripper: is the operator fighting the grasp?
Gripper metrics read the gripper dimension (averaged into one signal if you picked several), and they need the open direction you declared.
Metric
Measures
Read it
gripper_flips_per_s
Open/close transitions per second, with hysteresis so noise near the threshold doesn’t count
High means chatter or a hesitant operator.
missed_grasp_frac (opt-in)
Fraction of close commands where the follower closes fully with no stall residual, meaning nothing was held
High means empty grasps. Heuristic.
recovery_count (never scored)
Number of detected regrasps: open, then close again near the same place
Neutral or positive. Writes a recovery tag
Gripper metrics for chatter, empty grasps, and regrasp recovery.
Why it matters. A gripper that flips ten times in four seconds is either a noisy sensor or an operator who can’t get a grip. Both are worth a look.
Figure 5. Hysteresis keeps noise near the threshold from counting as a transition. Chatter scores 0.88 flips per second against 0.25 for a clean grasp. The regrasp has more transitions than the clean grasp and is tagged as a recovery instead of being penalized. Synthetic signals.
The design choice worth knowing. Several clean open-close cycles are often a recovery, an operator correcting a failed grasp. That’s behavior you may want a policy to learn. So recovery_count is never penalized, and a pause next to a regrasp isn’t counted against the time metrics either.
Pitfall. The gripper and tracking metrics have no published downstream validation. missed_grasp_frac and recovery_count are covered by synthetic tests only.
6. Consistency: does this demonstration agree with its peers?
action_divergence: same situation, different action
Measures. For frames in this episode, find the nearest state neighbors in other episodes of the same task group, and measure the variance of their action.
Intuition. If five demonstrations are in nearly the same state and four push left while one pushes right, the policy sees contradictory labels for the same input. Belkhale et al. (NeurIPS 2023) call this action divergence and report that “making actions more consistent tends to increase policy success.”
Use it when. You have enough episodes per task group for neighbors to exist (about 20).
Read it. High means this demonstration takes a different action from peers in a similar state.
Pitfall. Similar states can legitimately need different actions. A task with two valid strategies produces divergence that isn’t a defect. Treat it as a review signal.
Figure 6. Each arrow is the action another episode took from a state near this episode's. Variance of those actions: 0.00 on the left, 0.59 on the right. Schematic, in a 2D stand-in for the state space.
stats_leverage (never scored): how much one episode bends the normalization
Measures. How many feature dimensions this episode alone stretches beyond the rest of the dataset’s meta/stats.json bounds.
Why it matters. Training normally normalizes with stats.json. One episode with a sensor glitch that reports a joint value of 400 where everything else lives in -1 to 1 distorts that normalization for every other episode, and you’d never see it in a single-episode viewer.
7. Camera: can you trust the pixels?
Camera metrics are opt-in because they decode video. They’re pixel and signal arithmetic on a few decoded frames, with no model, so there’s nothing to download.
We store every value per camera and compare it with the same camera in other episodes. A wrist camera is never judged against a top camera. An episode’s camera score is its worst camera.
Metric
Measures
Read it
blur
The 10th percentile of Laplacian variance over 12 sampled frames, at a fixed size (longer side 480 px), reported as log10(1 + variance)
Lower is blurrier. The log scale exists because raw variance is heavy-tailed.
exposure_err
Distance of mean luminance from mid-gray, from 0 (mid-gray) to 1 (black or white). Median over sampled frames
High means the image is too dark or too bright.
clipped_frac
Share of pixels at pure black or pure white
High means blown highlights or crushed shadows.
frozen_frac
Of 6 short bursts, the share where frames are nearly identical while the state shows the robot moving
High means a frozen or dropped feed.
video_action_lag_ms
Absolute offset between motion in the video and arm speed in the action, from cross-correlation over the central 30 s
Near 0 is in sync.
Camera metrics, computed from a few decoded frames and compared with the same camera in other episodes.
Figure 7. The frame metrics on a synthetic scene, with values computed at the tile's resolution. The plugin measures at a fixed size (longer side 480 px) and compares each camera with the same camera in other episodes. A live camera on a static scene still shows frame-to-frame change (about 0.8 gray levels here). A frozen feed shows none.
Why frozen_frac is built the way it is. A still scene with a still robot isn’t a frozen camera. So a burst counts only when the robot is moving, and the pixels aren’t. A live camera on a static scene reads about 0.3 gray levels of frame-to-frame change or more, and the threshold (0.05) sits far below that.
Why video_action_lag_ms is the one to take seriously. If your video runs 400 ms behind your actions, the policy is trained to predict the action from an observation that’s already stale. It’s the failure that integrity checks can’t see, since durations match and every file is intact. The metric is relative to peers: a lag shared by every episode sets the baseline, and only episodes that differ from it stand out. Nothing is reported when the video doesn’t track the action (correlation under 0.3), for example when the robot is out of view.
Figure 8. Video motion energy trails the arm speed by 400 ms here. Cross-correlation finds the offset. Across episodes of one camera, the lag everyone shares is the baseline, and the episodes that differ from it stand out. Synthetic signals.
Use them when. You merge data from several sessions, several rigs, or the Hub. A camera problem that hits one session or one rig is easy to miss when you sample a few episodes, and obvious once you rank them all.
Pitfalls.
A scene that’s dark by design reads as high exposure_err, and black borders count as clipped. The per-camera comparison absorbs both, but a view that pools many robots will flag more episodes than a single-task dataset.
Cameras stored as images rather than video can’t be read, and aren’t offered.
Cost. Frame metrics share one decode of 12 frames, so enabling all three costs the same as enabling one. On the development set, a video takes about 0.05 s for the frame metrics and about 0.9 s per episode (all cameras) for the lag. Budget about 0.4 s per camera per episode when you run it.
8. Language: can a VLA learn from this instruction?
For a VLA, the task string is the conditioning input. It’s what the model sees alongside the images. A clean trajectory with the instruction "Hold" teaches the model very little about language.
Metric
Measures
task_missing
Empty, whitespace-only, or null task string
task_generic
Fewer than 3 tokens, a placeholder ("task desc", "Hold"), or no action verb
The two language checks run on each episode’s task string.
Task strings are canonicalized (case, whitespace, and punctuation) before grouping, so "Pick up the cube." and "pick up the cube" land in one group.
Why it’s separate. Language flags never touch the score, the verdict, or the flag count. They get their own verdict and their own column. A VLA user reads it next to the score. A policy trained without language conditioning can ignore it.
Keep it honest. Whether generic instructions hurt VLA language grounding hasn’t been measured in anything we reviewed. task_generic rests on the Hugging Face community’s list of these strings as real defects, which they documented while curating data for SmolVLA. Static checks also can’t tell you that a perfectly grammatical instruction describes the wrong video. That’s a model-based check, covered below.
9. Outliers: unusual is not the same as bad
Outlier metrics are never scored.
Metric
Measures
iforest_score
Isolation-forest anomaly score over the episode’s metric vector, fit within its task group
novelty_knn
Mean distance to the nearest same-group episodes in metric-vector space
is_outlier
Either score at z ≥ 2
The three outlier metrics, fit within each task group and never scored.
Use them for. A second look at episodes that no single metric flags but that sit far from everything else in the group.
Read them.novelty_knn is two-sided. Very high means unlike the rest. Very low means near-duplicate.
Why they never score. The length-confound audit (arXiv:2606.10229) found that isolation-forest curation matched no curation at all, and that outlier detectors can rank defective episodes as normal. It’s a single-author preprint on 80 scripted demonstrations, so read it as a caution rather than a verdict. The practical takeaway is this: a “weird” episode might be your most valuable edge case.
Every curation metric in this guide: what it reads, which direction is worse, whether it enters the score, and when to use it.
Beyond heuristics: the model-based metrics
Everything above runs without a model. That’s useful: nothing to download, nothing to host, and cheap, repeatable results. It also means certain questions stay unanswered. A heuristic can’t tell you that an instruction describes the wrong video, that two episodes are near-duplicates in what they show, or whether the robot actually completed the task.
The research literature has methods for each. They cost more, and their evidence is a mix of strong and thin.
One thing to be clear about: the plugin this post demonstrates computes none of these. It’s deliberately heuristic-only. FiftyOne provides the infrastructure for the model-based tier. It has embeddings and similarity through the FiftyOne Brain, a model zoo (including remote zoo models) for running vision-language models (VLMs), and the same panel and tagging workflow to review whatever those produce. You’d build these as additional operators on the same pattern.
1. Instruction-video agreement (VLM). Sample frames, ask a vision-language model whether the video shows the instruction, and store the answer, a confidence score, and a proposed rewrite. SmolVLA’s data pipeline did a version of this, using Qwen2.5-VL-3B to rewrite task strings from sample frames plus the original label, with a prompt asking for under 30 characters starting with an action verb (“Pick,” “Place,” “Open”). Watch out: VLM judges are noisy, and OpenGVL’s benchmark found open VLMs lag proprietary ones. Keep the result as its own verdict, never folded into a score.
2. Embedding-based duplicates and novelty. Embed sampled frames and the action trajectory, then find near-duplicates by cosine similarity. SCIZOR does this at the transition level, using a joint state-action embedding, and reports that its two thresholds (εs = 0.58, εd = 0.99) transferred across RoboMimic, Open X-Embodiment, and real data. Watch out: near-duplicates in a table-top task might be exactly the repetition that robust policies need. Duplication needs a human decision about how much is too much.
3. Task progress. Ask a model to order shuffled frames by progress, then measure how well that order matches the true frame order (Value-Order Correlation, from GVL). OpenGVL applied this to LeRobot Hub data and found task-definition problems, labeling ambiguity, and failed or out-of-distribution episodes. TOPReward replaces generated numbers with token probabilities and reports VOC of 0.874 with Qwen3-VL-8B against 0.218 for GVL with the same model on a filtered Open X-Embodiment subset. Watch out: these are model-dependent, and the numbers come from the authors’ own benchmarks.
4. Information-based scoring. DemInf (Hejna et al., RSS 2025) estimates each demonstration’s contribution to state-action mutual information with k-nearest neighbor (kNN) estimators over learned embeddings. Its abstract reports a 5-10% improvement on RoboMimic and better performance on real ALOHA and Franka setups. The authors note it works best with relative actions.
5. Policy-in-the-loop. CUPID uses influence functions to estimate each demonstration’s effect on expected return from evaluation rollouts, and reports that under 33% of curated data yields state-of-the-art RoboMimic diffusion policies. Demo-SCORE trains a classifier on successful versus failed policy rollouts and filters demonstrations that resemble the failures, and its abstract reports over 15-35% higher absolute success. These need a policy and rollouts, so they’re the highest-fidelity and the most expensive tier.
The CUPID finding to remember. In CUPID’s results, DemInf curated the “highest overall quality,” yet CUPID-curated policies matched or beat it. The authors write that “human perception of demonstration quality does not necessarily correspond to data that maximizes downstream policy success.” That’s a direct warning about every heuristic score in this guide, including the ones that look best.
Dataset-level metrics
Some questions aren’t about any one episode.
Task coverage. Episodes per canonical task. A dataset with 1,900 episodes of one task and 12 of another isn’t balanced, no matter what the total says.
Source balance. Episodes per source repo when you’ve merged several.
Mixture weights. If you’re combining datasets, the weights matter as much as the filtering. Re-Mix (Hejna et al., CoRL 2024) optimizes domain weights and reports beating uniform weights by 38%. Octo and OpenVLA both curated mixtures from Open X-Embodiment by hand. Octo dropped datasets without images or delta end-effector control, plus ones “too repetitive, low image resolution, or excessively niche.” OpenVLA kept single-arm, third-person-camera, end-effector datasets.
The panel reports coverage as a count on the Integrity & Coverage tab. It’s a count, not a pick, since the right balance depends on what you’re training for.
From metrics to a ranking
Dozens of metrics don’t help you decide where to look. One ranking does. Getting there takes two decisions: compared to what, and how to combine them.
Compared to what: normalization
Task strings are canonicalized, then episodes are grouped by task (an episode with several tasks is grouped by its sorted set). Groups with at least 20 episodes are scored against themselves. A smaller group is pooled across views, and a view with fewer than 20 episodes overall is pooled and marked low-confidence in the panel.
The scoring uses robust z-scores with a floored scale. Zero-inflated metrics (such as idle_lead_s and frame_gaps, where most episodes score exactly zero) use percentile ranks or a tail-based scale, because a standard deviation computed over mostly zeros is meaningless.
Why per task. Without it, you rank tasks against each other. When we ran an earlier version of this scoring on a multi-task robot dataset, two tasks topped the ranking as a group. That’s the cross-task effect: a batch that mixes tasks produces misleading outliers, because “normal” differs by task. Score episodes of the same task together.
Figure 9. A slower, smaller task (B) pooled with a larger one (A). Pooled, the warn line sits near A's distribution and flags 26 of B's 30 episodes for being task B. Per task, each group is flagged against its own distribution: 2 of 120 in A and 1 of 30 in B. Synthetic data, robust z with a median and MAD scale.
Combined how: groups, and the max
Correlated metrics share a group, and a group counts as one vote:
The scored members of each metric group. Each group counts as one vote in the episode score.
Inside a group, the value is the maximum of the weighted, oriented member z-scores. Across groups, the episode’s score is the maximum group value, with the weighted mean as a tie-breaker. Weights are 1.0 except ldlj, idle_trail_s, and end_motion_ratio at 0.5.
The effect is that five fine ones can’t average away one severe problem in one dimension. An episode with a frozen camera and perfect motion still ranks near the top. Each episode also carries n_flags (how many groups are flagged) and driver (which group is behind its score), so you know why it ranks where it does.
Figure 10. An example episode with made-up z-scores. Every group is fine except camera, where frozen_frac is at 4.3. The score is the maximum group value, n_flags counts groups at or above the warn line (z = 2), and driver names the group behind the score.
The workflow, step by step
A metric earns its place by changing what you do next. Here’s the loop.
Step 1: Import the dataset
One sample is one episode. Set VFF_MULTIMODAL=1 before importing FiftyOne in every process that touches a LeRobot dataset, since it enables the multimodal episode viewer. Install the plugin with the plugin manager or link a checkout into your plugins folder.
import fiftyone as fo
dataset = fo.Dataset.from_dir(
dataset_dir="/data/lerobot/my_dataset",
dataset_type=fo.types.LeRobotDataset,
)
# Merge another source if you want to curate across datasets
dataset.add_dir(
dataset_dir="/data/lerobot/other_dataset",
dataset_type=fo.types.LeRobotDataset,
)
session = fo.launch_app(dataset)
See the App guide for sidebar filters and Spaces, which is where the curation panel opens.
Step 2: Compute quality
Open the operator browser and run LeRobot curation: compute quality. It’s delegated by default, so it runs in the background. The form has four tabs: Data (what you tell the tool), Metrics (a checkbox per metric, grouped by family), Camera, and Normalization (the smallest task group scored on its own, default 20).
Your picks are remembered, so re-running doesn’t start from scratch.
Step 3: Read the Overview
Open a panel and choose LeRobot Curation. The Overview tab has the score histogram, verdict counts for Integrity and Language, episodes per task, the outlier scatter, and the worst-first ranking. The panel follows the view: filter the grid and it re-ranks.
Step 4: Drill into the family that drives the problem
Use the driver column to pick a tab. Motion & Action shows one histogram per metric. Vision shows the five camera metrics with camera chips. Click any bar and the samples grid filters to those episodes.
Step 5: Open the inspector
Click any row. The inspector shows every metric with its per-arm or per-camera breakdown, joint traces (action solid, state dashed), the speed profile with the idle threshold, the gripper timeline, and a few frames from the cameras, picked around the flagged spans.
Flagged spans (idle stretches, the longest pause, the roughest smoothness window, and acceleration spikes) are also written as temporal tags on the episode’s timeline once their metric reaches warn so that you can scrub straight to them. Regrasp recoveries are tagged too, as information. A re-run replaces the plugin’s tags and leaves any tag you drew yourself alone.
Step 6: Tag, don’t delete
The footer has three tag buttons: review, exclude-candidate, and relabel. They tag the selection, or everything in view if nothing is selected. Nothing is deleted or hidden. These are ordinary FiftyOne sample tags, so you can filter on them with view stages, save a view, or act on them from code.
Marking an episode exclude-candidate is always a separate human action. The scores never do it for you.
The native exporter (API reference) writes a self-contained LeRobot v3 dataset. export_media=True is required, episodes are renumbered contiguously, and task and frame indexes are rebuilt. From there, you can push it to the Hub as you would any LeRobot dataset.
If you’d rather edit in place, send the excluded episode_index values to lerobot-edit-dataset with the delete_episodes operation. Task string rewrites are a separate matter: as of the sources we reviewed, lerobot-edit-dataset can’t yet edit task strings, so accepted rewrites need to be patched into meta/tasks.parquet and the per-frame task_index after export.
Keep it honest
Scores are triage, not verdicts. Smoothness and timing say nothing about whether the demonstration did the right thing. The plugin’s own documentation puts it this way: use it to decide where to look first, and don’t use it as an automatic accept/reject gate.
The evidence is genuinely mixed.
Smoothness filtering helped in RINSE, but not in behavior cloning with proprioceptive signals on episodes already filtered for success.
The curation-metrics audit found 5 of 7 metrics exploited episode length, and detection accuracy didn’t predict policy success. That’s one preprint on 80 scripted demonstrations. It’s also exactly why outlier metrics aren’t scored and why the length check exists.
CUPID found that data humans would call high quality can underperform data selected by influence on the policy.
The gripper and tracking metrics have no published downstream validation.
The thresholds are starting points. Defaults such as the 2% joint-limit margin, the warn line at z ≥ 2, the 20-episode group size, and the camera thresholds were calibrated on a 102-episode development set. They aren’t yet adjustable in the form. Check them on your own data.
How the metrics are checked. Each scored metric has a corruption it must catch: added hesitation, an inserted idle run, a duplicated frame, a blanked task, a stalled joint, and a blurred or frozen video. The harness applies the corruption to a real episode and asserts that the corrupted copy ranks worse than the original. It also re-scores every episode truncated to a common length, to find metrics that only measure duration. On 102 episodes, it takes about 5 minutes with camera metrics included. That tells you a metric detects what it claims to detect. It can’t tell you the detection predicts policy quality.
So validate. Before you drop data, train on the curated subset and the full set, then compare. A curation run you haven’t validated with a training run is a hypothesis.
Where FiftyOne fits
Rerun and Foxglove are excellent at what they’re built for: looking closely at a recording. If you want a synchronized view of cameras, joint traces, and 3D for one episode, they’re good tools.
Curation asks a different question, which is about a collection. Which of these 10,000 episodes should I look at first? Of those, which ones have a camera problem? Show me only the ones from this session. Tag these 40. Give me everything else as a dataset I can train on.
That needs a dataset model rather than a viewer: episodes as queryable samples, fields you can filter on, charts that filter the grid when you click them, tags that persist, saved views, and an export that round-trips to the training format. FiftyOne is built around that model, and that’s why the workflow above works the way it does.
The part that matters most if you build your own is how little of this is fixed. FiftyOne’s plugin framework is Python operators plus React panels, and you can mix the two in hybrid panels. The panel in this guide is a React app backed by operators. One data operator returns the whole payload (rows, histogram bins, thresholds, and verdict counts) and re-fires whenever the view changes. A second operator returns a single episode’s arrays and frames on demand for the inspector. The five tabs are handwritten React. Adding a metric is one Python function and one dictionary entry, and charting it means copying a documented chart template.
If your rig has a metric nobody else computes, such as force-torque spikes, a custom contact signal, or a per-operator score, you add it, and it gets the same filter, tag, inspector, and export workflow as everything else.
Start with the integrity checks (frame_gaps, length_mismatch, nonfinite_values, too_short), because they’re free and tell you whether the rest is worth running. Then add sparc for motion, idle_lead_s and longest_pause_s for dead time, and frozen_frac if you can afford video decode.
Both measure smoothness of the speed profile. SPARC measures how many corrections the motion contains, via its frequency spectrum, and is scale-invariant. It responds to stop-and-go motion but not to additive white noise. LDLJ is log dimensionless jerk. It does respond to white noise, but it’s noisier and duration-sensitive, so it carries half weight in the score.
Different tasks have different normal motion and timing. A pooled score ranks the slowest, most correction-heavy task as the problem. Per-task normalization compares each episode to its own peers. It needs about 20 episodes per task. With fewer, the panel pools the view and shows a low-confidence banner.
It’s mostly a hardware property. A slow servo loop delays every episode by roughly the same amount. The metric is useful for comparing rigs and sessions, so we summarize it per dataset and flag only episodes far from the median.
Nobody has measured it in anything we reviewed. The task_generic check rests on the Hugging Face community’s documented list of such strings as real defects. A VLA uses the instruction as conditioning input, so a weak string is a reasonable thing to review. It has its own verdict and never enters the score.
We recommend against it. The published evidence is mixed: smoothness filtering helped in one study, while another found most curation metrics exploited episode length and that detection accuracy didn’t predict policy success. Use the ranking to decide what to review, tag exclude-candidate yourself, and validate with a training run.
No. Every metric is signal or pixel arithmetic, so there’s nothing to download. Model-based checks such as VLM instruction verification and embedding-based duplicate detection are deliberately left out. FiftyOne provides the embeddings and similarity and model zoo infrastructure to build them as extra operators.
Filter out the exclude-candidate tag with dataset.match_tags("exclude-candidate", bool=False) and call .export(..., dataset_type=fo.types.LeRobotDataset, export_media=True). The exporter writes a self-contained v3 dataset with episodes renumbered contiguously.
Harpreet Sahota
Hacker-in-Residence
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.