You Can’t Watch 10,000 Episodes. Measure These Instead: Curation Metrics for LeRobot Datasets

Oct 8, 2026
•
27 min read
Author
Headshot of Harpreet Sahota
Harpreet Sahota
Hacker-in-Residence
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.
See all articles by Harpreet Sahota

Talk to an AI expert

A field guide to the metrics worth computing on a LeRobot dataset before you train a policy or fine-tune a vision-language-action (VLA) model: what each one measures, when to trust it, where it misleads you, and how to run the whole workflow in FiftyOne.
You recorded 400 teleop episodes last week. You pulled another 6,000 from the Hugging Face Hub because the task looked close to yours. Now you want to fine-tune a VLA.
You can’t watch all of it. At 30 seconds per episode, 10,000 episodes is about 83 hours of video, and that’s before you open a second camera. So you scrub through a handful, they look fine, and you hit train.
Then the policy hesitates at the first grasp. Somewhere in there is an episode with a frozen wrist camera. Another has 11 seconds of dead time before the arm moves. A third has the task string "Hold". A fourth has a video track that runs 0.4 seconds behind the actions. None of these show up when you watch one episode at a time because none of them is visible in a single episode. They show up when you compare an episode to its peers.
That’s the gap this post is about. A viewer shows you data. Curation tells you which episode to open next.
This guide walks through the metrics that do that job for LeRobot and VLA datasets. For each one, it covers what it measures, why you’d compute it, when to use it, how to read it, and where it lies to you. It also shows where each metric lives in a FiftyOne panel, because a metric you can’t filter, tag, and export is just a number in a notebook.
The post isn’t about any one dataset. The metrics apply to anything in LeRobot v3 format.

Key takeaways

  • Curation metrics answer six different questions: is the episode broken (integrity), is the motion clean (smoothness), is it wasting frames (time), did the robot do what was commanded (tracking), can you trust the cameras (vision), and can a model learn from the instruction (language). Start with the cheapest.
  • Most of these metrics need no model and no GPU. Integrity checks read metadata and Parquet. Motion, time, tracking, gripper, and consistency metrics are signal math on action and observation.state. Only the camera metrics decode video, at roughly 0.4 s per camera per episode.
  • Compare episodes to their own task group, not to the whole dataset. A fold-a-shirt episode and a pick-a-cube episode have different “normal” motion, and a pooled score will rank one task as the problem.
  • Correlated metrics should share one vote. Five time metrics that all fire on the same pause are one problem, not five.
  • The published evidence is mixed. Spectral arc length (SPARC) smoothness filtering raised RoboMimic success by 16% using one-sixth of the data (RINSE, arXiv:2604.23000). A separate audit (arXiv:2606.10229) found 5 of 7 curation metrics exploited episode length, and that detection accuracy and policy success were uncorrelated. Use scores to order a review queue, and validate with a training run before you delete anything.
  • FiftyOne’s role is the workflow around the metrics: one sample per episode, a panel that re-ranks as you filter, click-to-filter charts, temporal tags on the episode timeline, and a lossless export of the kept subset as a LeRobot v3 dataset.

First, the LeRobot v3 format, in one paragraph

A LeRobot v3 dataset doesn’t store one file per episode. Frames from many episodes are concatenated into shared Parquet files (data/chunk-000/file-000.parquet), and each camera’s video is concatenated into shared MP4s. Metadata in meta/info.json holds the schema and fps, meta/stats.json holds normalization statistics, and meta/episodes/ holds one row per episode with its row range and, per camera, the video time window.
Two consequences matter for curation. First, “episode N” points into shared files, so bookkeeping can drift without anything crashing. Second, the timestamp column is written from the nominal fps, not from a logged clock, so classic sensor-health checks such as jitter and clock drift don’t apply. You check structure and alignment instead.
FiftyOne’s native fo.types.LeRobotDataset (FiftyOne 1.22 and later) imports each episode as one multimodal sample, with a synchronized episode viewer on top (API reference). It reads only the metadata up front and pulls Parquet rows and video frames on demand, which is what makes dataset-level scoring practical.

What you have to tell the tool

Nothing about your robot should be guessed. A 14-dimensional action vector could be two 7-degree-of-freedom (DoF) arms, or one arm plus a hand, or end-effector deltas. If a tool silently picks, every smoothness score downstream is wrong in a way you can’t see.
So the compute form asks for five things:
  • Which arrays are state and action. They default to observation.state and action when the dataset uses the LeRobot standard names, and stay unset otherwise.
  • What the action means. Is it absolute joint positions in the same space as the state, or something else (deltas, end-effector poses)? This is a yes/no question, and it only appears when the two arrays have the same size.
  • Robot and joints. Single or dual arm, whether there’s a multi-joint hand, and which joints are the arm and which are the gripper. You pick them by name, with each joint’s observed range beside it.
  • Gripper open direction. Does a high value mean open or closed?
  • Cameras. Which video cameras to score. None are selected by default, because decoding is the slow part.
Anything you leave unanswered switches off the metrics that need it, and the form lists exactly which. It never blocks the run.
Good to know. Why does “what the action means” matter so much? Tracking, acceleration-spike, and joint-limit metrics compare the commanded action to the achieved state. That comparison only means something if both live in the same joint space. If your action is end-effector deltas, a gap between the two is just a unit mismatch, so those metrics stay off.

The LeRobot curation metrics

Every metric below follows the same pattern: what it measures, why it matters, when to use it, how to read it, and the pitfall to watch for. The harder ones get an intuition-level explanation.
One convention first. Inside a score, every metric is oriented so that higher means worse. The tables and charts show each metric’s raw direction, and the scoring section explains how they combine.

1. Integrity: is this episode broken?

Integrity checks don’t say an episode is worse. They say it’s broken. They produce pass, warn, or fail flags and counts. They never enter the ranking score, and they get their own verdict. They’re also the cheapest checks you’ll run, since they read metadata and Parquet and decode nothing.
Run them first. There’s no point ranking the smoothness of an episode whose rows don’t match its video.
Figure 1. A v3 episode is a row range in a shared Parquet file plus a time window in a shared MP4 per camera. Each integrity check catches a case where a pointer stopped matching the data. Illustrative values.
CheckWhat it catches
frame_gapsNon-contiguous or duplicated frame_index values. Any nonzero count means a broken episode.
length_mismatchThe episode’s declared length disagrees with the actual number of Parquet rows. Metadata and data disagree.
video_window_mismatchPer camera, the video window duration disagrees with length/fps. Video and actions differ in length.
nonfinite_valuesNaN or inf in action or state. Any nonzero count breaks training.
too_shortEpisodes under a minimum duration or frame count. Catches one-frame and near-empty episodes.
schema_mismatchWhen you merge sources: differences in feature shapes and names, fps, units, camera keys, or resolution. Flags a dataset that won’t collate cleanly.
The six integrity checks and the broken-episode case each one catches.
timestamp_dev (the largest deviation of consecutive timestamps from 1/fps) is reported as information only. Because LeRobot writes timestamps from fps, it’s usually about zero and doesn’t affect the verdict.
Why it matters. The Hugging Face team’s write-up on community datasets lists these defects as recurring: episodes with one or very few frames, Parquet files deleted without reindexing, and inconsistent action or state dimensions for the same robot. Each one fails silently until a dataloader hits it, or worse, doesn’t fail at all.
Pitfall. video_window_mismatch compares durations, so it can’t see a constant offset. An episode whose video starts 400 ms late passes. That’s what the camera lag metric later in this guide is for.

2. Motion: is the motion clean?

Motion metrics measure the shape of the movement an operator commanded. This is where most of the curation literature lives, and it’s also where the most care is needed, so the preprocessing matters as much as the metric.
What signal it reads. The action array over the joints you put in the Arm (and Hand) fields. action is what a policy learns to output. On leader-follower rigs such as SO-100, Koch, and ALOHA, it’s the operator’s hand, not the robot’s arm.
How the signal is prepared.
  • Each joint is range-normalized before speeds are combined, so a joint with a large range doesn’t dominate and units don’t matter.
  • Gripper and hand dimensions are excluded from arm speed. A 22-joint dexterous hand is scored as its own signal so it never swamps a 7-joint arm.
  • The low-pass cutoff is min(10 Hz, 0.4 × fps). At 30 fps that’s 10 Hz, well away from the Nyquist frequency of 15 Hz.
Metrics run on fixed-length windows because some are sensitive to episode duration. Each scored metric also keeps a worst-window value, so one bad stretch isn’t averaged into a fine episode.

sparc: how many corrections does the motion contain?

Measures. The spectral arc length of the arm’s speed profile.
Intuition. Take the arm's speed over time. A clean reach is one smooth bump: speed rises, peaks, falls. Take that curve’s frequency spectrum and trace along it. A single bump has a compact spectrum, so the traced path is short. Every extra correction, hesitation, or stop-and-go adds frequency content, and the path gets longer. SPARC is the negative of that path length, so smooth motion sits near 0, and fragmented motion goes very negative.
In its standard form, SPARC is the negative arc length of the normalized speed spectrum up to a cutoff frequency:
SPARC = -∫ sqrt( (1/ω_c)² + (dV̂(ω)/dω)² ) dω (0 to ω_c)
where V̂(ω) is the magnitude spectrum of the speed profile, normalized by its peak. You don’t need to compute that by hand to use the metric. The picture matters: more corrections mean a longer trace.
Why it matters. It’s the most direct validation in the literature. RINSE (arXiv:2604.23000) reports that SPARC-based filtering gave 16% higher RoboMimic success using one-sixth of the data. It also reports SPARC as about 10 times more noise-robust than log dimensionless jerk (LDLJ).
Use it when. You want one primary smoothness number per episode. It’s scale-invariant, so it works across arms of different sizes.
Read it. Closer to 0 is smoother. Very negative means fragmented, hesitant motion.
Pitfall. SPARC responds to stop-and-go motion, not to additive white noise. A jittery signal from a noisy encoder can still look smooth to SPARC. That’s what the next metric is for.

ldlj: how much jerk, once you account for size and speed?

Measures. Log dimensionless jerk of the speed profile. Jerk is the rate of change of acceleration. It’s integrated over the movement, made dimensionless using the movement’s duration and peak speed, then log-transformed.
Why it matters. It catches what SPARC doesn’t: high-frequency noise and abrupt changes.
Use it when. You suspect noisy teleop hardware or a bad leader-arm encoder, alongside sparc.
Read it. Closer to 0 is smoother.
Pitfall. It’s noisier than SPARC and duration-sensitive, which is why the windows are fixed-length and why it carries half weight in the score.
Figure 2. The same two reaches, moved three ways. SPARC: smooth -2.57, stop-and-go -5.25, jittery -2.56. LDLJ: smooth -8.87, stop-and-go -14.45, jittery -14.71. SPARC separates stop-and-go from the other two. LDLJ penalizes both stop-and-go and noise. Synthetic signals at 30 fps. Values computed from them.

sparc_phase: smoothness within each phase of the grasp (opt-in)

Measures. SPARC computed inside each gripper-delimited phase (approach, grasp, transport, and release), rolled up by median.
Why it matters. Contact is inherently jerky. When the gripper closes on an object, speed changes abruptly for good reasons. Scoring the whole episode as one signal counts those transitions as roughness. Scoring per phase stops that.
Use it when. Your task has clear grasp and release events, and you’ve declared gripper joints and the open direction.
Read it. Same as sparc.
Pitfall. It adapts the gripper-phase idea from RINSE and isn’t the paper’s exact method. It’s covered by synthetic tests only. It skips phases with too few samples.
Two more motion metrics exist for review only and are never scored. jerk_rms is the root mean square (RMS) jerk after a low-pass filter. It’s noise-sensitive and correlated with the other metrics. psd_lf_hf is the log ratio of low- to high-frequency power in the speed profile, and it’s unreliable at LeRobot frame rates, where tremor sits near Nyquist.
Keep it honest. RINSE assumed its episodes were already filtered for task success, and it tested behavior cloning with proprioceptive signals, not VLAs. Smoothness helps. It isn’t sufficient. A smooth episode can still show the wrong task.

3. Time efficiency: is the episode wasting frames?

Time metrics measure dead time. They read the commanded motion (action), with an idle threshold set as a multiple of the episode’s own moving speed. That per-episode threshold means a slow, careful operator isn’t flagged just for being slow.
MetricMeasuresRead it
idle_lead_sSeconds of stillness at the startHigh means dead time before the task begins. Also stores a suggested trim start (keep_from_s).
idle_trail_sSeconds of stillness at the endHigh means a long dead tail. Stores a suggested trim end (keep_to_s).
longest_pause_sLongest idle run that isn’t at either endHigh means a mid-task stall or hesitation.
end_motion_ratioSpeed in the last second relative to typical moving speedHigh means the episode probably ended mid-motion.
length_zRobust z of episode duration within its task groupScored as abs(z). Positive means unusually long (struggle). Negative means unusually short (abandoned).
Time efficiency metrics, what each one measures, and how to read a high value.
Figure 3. Left: one episode's arm speed with the three idle stretches the time metrics measure, and the keep window they suggest. Right: length_z places an episode's duration against its task group. Synthetic signals.
Why it matters. Two published pipelines treat this as signal. OpenVLA removed all-zero no-op actions from its training data. SCIZOR’s removed samples included “manipulation failures, slow motion, and pauses.” Both pipelines treated dead frames as removable, and either way they cost training compute.
Use it when. Your data comes from human teleop, where operators pause to think, reposition, or wait for a timer.
Pitfalls.
  • Trailing idle and end_motion_ratio get half weight, because fixed-duration recording setups produce a dead tail from the timer itself.
  • A pause next to a regrasp counts as recovery and is excluded from longest_pause_s. A hesitation that precedes a retry is the operator fixing a problem, not wasting time.
  • idle_frac (the overall fraction of idle frames) is shown for context and never scored. A high value can be a legitimate pause. path_length_z (opt-in) is also never scored, since it correlates with length_z.
  • length_z, path_length_z, and idle_frac depend on episode length by construction. Only length_z is scored, and it shares one vote with the idle metrics. That design decision comes straight from the length-confound audit discussed later.
Trimming is a suggestion. keep_from_s and keep_to_s store a proposed range. Nothing is written back to the dataset.

4. Tracking and contact: did the robot do what was commanded?

On a leader-follower rig, action is where the operator’s leader arm is, and observation.state is where the follower robot actually is. When the two agree, the robot is doing what it was told. When they diverge, something physical happened: a collision, a stall, a slipping joint, or a following arm that couldn’t keep up.
These metrics read the state, so they run only when you’ve declared the action to be absolute joint positions in the state space.
Figure 4. The follower trails the leader by a few frames (right), then stalls while the command keeps moving. The residual after lag alignment lights up exactly where the stall is, and the catch-up produces an acceleration spike. Synthetic signals.

track_resid: how far is the follower from the command?

Measures. The per-joint, normalized gap between action and state after lag alignment, with each joint’s median offset removed first. Arm joints only.
Why it matters. Calibration offsets are constant and harmless. You want the part of the gap that varies, where contact and slippage live.
Read it. High means the robot isn’t reaching the commanded position. It also raises interval flags that become spans on the timeline.

track_lag_ms: how late is the follower?

Measures. The delay between command and response, estimated by sliding one signal against the other and picking the offset with the highest cross-correlation:
lag = argmax_k corr( action[t], state[t + k] )
Why it matters. Mostly it’s a hardware property. A servo with a slow control loop lags every episode by about the same amount.
Read it. It’s summarized per dataset or session, and only episodes far from the dataset median are flagged. It never enters a score.

accel_spike_frac: sudden jolts

Measures. The fraction of frames where any joint’s state acceleration exceeds median + k × MAD (median absolute deviation), with an absolute floor.
Why it matters. It’s a proxy for collisions and contact. The absolute floor stops idle joints from turning every frame into a spike.

joint_limit_frac (opt-in): operating near the edge

Measures. The share of (frame, joint) pairs within a margin (default 2% of range) of a joint limit.
Pitfall. It uses your limits if you give them, and otherwise falls back to meta/stats.json min and max. That’s a weak proxy, because observed extremes always sit at the edge of the observed range.
Keep it honest. These metrics have no published downstream validation. Treat them as sanity checks. Related work (RoboDrop, arXiv:2609.10021) shows observation-action temporal misalignment is a major corruption that action-only methods handle poorly, and a teleoperation paper (arXiv:2605.26349) flags stalls and operation near joint limits, but its user study had 3 operators. That’s weak evidence, not a validated benchmark.

5. Gripper: is the operator fighting the grasp?

Gripper metrics read the gripper dimension (averaged into one signal if you picked several), and they need the open direction you declared.
MetricMeasuresRead it
gripper_flips_per_sOpen/close transitions per second, with hysteresis so noise near the threshold doesn’t countHigh means chatter or a hesitant operator.
missed_grasp_frac (opt-in)Fraction of close commands where the follower closes fully with no stall residual, meaning nothing was heldHigh means empty grasps. Heuristic.
recovery_count (never scored)Number of detected regrasps: open, then close again near the same placeNeutral or positive. Writes a recovery tag
Gripper metrics for chatter, empty grasps, and regrasp recovery.
Why it matters. A gripper that flips ten times in four seconds is either a noisy sensor or an operator who can’t get a grip. Both are worth a look.
Figure 5. Hysteresis keeps noise near the threshold from counting as a transition. Chatter scores 0.88 flips per second against 0.25 for a clean grasp. The regrasp has more transitions than the clean grasp and is tagged as a recovery instead of being penalized. Synthetic signals.
The design choice worth knowing. Several clean open-close cycles are often a recovery, an operator correcting a failed grasp. That’s behavior you may want a policy to learn. So recovery_count is never penalized, and a pause next to a regrasp isn’t counted against the time metrics either.
Pitfall. The gripper and tracking metrics have no published downstream validation. missed_grasp_frac and recovery_count are covered by synthetic tests only.

6. Consistency: does this demonstration agree with its peers?

action_divergence: same situation, different action

Measures. For frames in this episode, find the nearest state neighbors in other episodes of the same task group, and measure the variance of their action.
Intuition. If five demonstrations are in nearly the same state and four push left while one pushes right, the policy sees contradictory labels for the same input. Belkhale et al. (NeurIPS 2023) call this action divergence and report that “making actions more consistent tends to increase policy success.”
Use it when. You have enough episodes per task group for neighbors to exist (about 20).
Read it. High means this demonstration takes a different action from peers in a similar state.
Pitfall. Similar states can legitimately need different actions. A task with two valid strategies produces divergence that isn’t a defect. Treat it as a review signal.
Figure 6. Each arrow is the action another episode took from a state near this episode's. Variance of those actions: 0.00 on the left, 0.59 on the right. Schematic, in a 2D stand-in for the state space.

stats_leverage (never scored): how much one episode bends the normalization

Measures. How many feature dimensions this episode alone stretches beyond the rest of the dataset’s meta/stats.json bounds.
Why it matters. Training normally normalizes with stats.json. One episode with a sensor glitch that reports a joint value of 400 where everything else lives in -1 to 1 distorts that normalization for every other episode, and you’d never see it in a single-episode viewer.

7. Camera: can you trust the pixels?

Camera metrics are opt-in because they decode video. They’re pixel and signal arithmetic on a few decoded frames, with no model, so there’s nothing to download.
We store every value per camera and compare it with the same camera in other episodes. A wrist camera is never judged against a top camera. An episode’s camera score is its worst camera.
MetricMeasuresRead it
blurThe 10th percentile of Laplacian variance over 12 sampled frames, at a fixed size (longer side 480 px), reported as log10(1 + variance)Lower is blurrier. The log scale exists because raw variance is heavy-tailed.
exposure_errDistance of mean luminance from mid-gray, from 0 (mid-gray) to 1 (black or white). Median over sampled framesHigh means the image is too dark or too bright.
clipped_fracShare of pixels at pure black or pure whiteHigh means blown highlights or crushed shadows.
frozen_fracOf 6 short bursts, the share where frames are nearly identical while the state shows the robot movingHigh means a frozen or dropped feed.
video_action_lag_msAbsolute offset between motion in the video and arm speed in the action, from cross-correlation over the central 30 sNear 0 is in sync.
Camera metrics, computed from a few decoded frames and compared with the same camera in other episodes.
Figure 7. The frame metrics on a synthetic scene, with values computed at the tile's resolution. The plugin measures at a fixed size (longer side 480 px) and compares each camera with the same camera in other episodes. A live camera on a static scene still shows frame-to-frame change (about 0.8 gray levels here). A frozen feed shows none.
Why frozen_frac is built the way it is. A still scene with a still robot isn’t a frozen camera. So a burst counts only when the robot is moving, and the pixels aren’t. A live camera on a static scene reads about 0.3 gray levels of frame-to-frame change or more, and the threshold (0.05) sits far below that.
Why video_action_lag_ms is the one to take seriously. If your video runs 400 ms behind your actions, the policy is trained to predict the action from an observation that’s already stale. It’s the failure that integrity checks can’t see, since durations match and every file is intact. The metric is relative to peers: a lag shared by every episode sets the baseline, and only episodes that differ from it stand out. Nothing is reported when the video doesn’t track the action (correlation under 0.3), for example when the robot is out of view.
Figure 8. Video motion energy trails the arm speed by 400 ms here. Cross-correlation finds the offset. Across episodes of one camera, the lag everyone shares is the baseline, and the episodes that differ from it stand out. Synthetic signals.
Use them when. You merge data from several sessions, several rigs, or the Hub. A camera problem that hits one session or one rig is easy to miss when you sample a few episodes, and obvious once you rank them all.
Pitfalls.
  • A scene that’s dark by design reads as high exposure_err, and black borders count as clipped. The per-camera comparison absorbs both, but a view that pools many robots will flag more episodes than a single-task dataset.
  • Cameras stored as images rather than video can’t be read, and aren’t offered.
Cost. Frame metrics share one decode of 12 frames, so enabling all three costs the same as enabling one. On the development set, a video takes about 0.05 s for the frame metrics and about 0.9 s per episode (all cameras) for the lag. Budget about 0.4 s per camera per episode when you run it.

8. Language: can a VLA learn from this instruction?

For a VLA, the task string is the conditioning input. It’s what the model sees alongside the images. A clean trajectory with the instruction "Hold" teaches the model very little about language.
MetricMeasures
task_missingEmpty, whitespace-only, or null task string
task_genericFewer than 3 tokens, a placeholder ("task desc", "Hold"), or no action verb
The two language checks run on each episode’s task string.
Task strings are canonicalized (case, whitespace, and punctuation) before grouping, so "Pick up the cube." and "pick up the cube" land in one group.
Why it’s separate. Language flags never touch the score, the verdict, or the flag count. They get their own verdict and their own column. A VLA user reads it next to the score. A policy trained without language conditioning can ignore it.
Keep it honest. Whether generic instructions hurt VLA language grounding hasn’t been measured in anything we reviewed. task_generic rests on the Hugging Face community’s list of these strings as real defects, which they documented while curating data for SmolVLA. Static checks also can’t tell you that a perfectly grammatical instruction describes the wrong video. That’s a model-based check, covered below.

9. Outliers: unusual is not the same as bad

Outlier metrics are never scored.
MetricMeasures
iforest_scoreIsolation-forest anomaly score over the episode’s metric vector, fit within its task group
novelty_knnMean distance to the nearest same-group episodes in metric-vector space
is_outlierEither score at z ≥ 2
The three outlier metrics, fit within each task group and never scored.
Use them for. A second look at episodes that no single metric flags but that sit far from everything else in the group.
Read them. novelty_knn is two-sided. Very high means unlike the rest. Very low means near-duplicate.
Why they never score. The length-confound audit (arXiv:2606.10229) found that isolation-forest curation matched no curation at all, and that outlier detectors can rank defective episodes as normal. It’s a single-author preprint on 80 scripted demonstrations, so read it as a caution rather than a verdict. The practical takeaway is this: a “weird” episode might be your most valuable edge case.

Master reference table

MetricReadsDirectionIn score?Use it when
frame_gaps, length_mismatch, nonfinite_values, too_shortmetadata, Parquetany nonzero is badOwn verdictAlways, first
video_window_mismatch, schema_mismatchmetadata, video windowsany mismatch is badOwn verdictMerging sources, any camera data
sparcactioncloser to 0 is smootherYesAlways
ldljactioncloser to 0 is smootherYes, half weightNoisy hardware
sparc_phaseaction + grippercloser to 0 is smootherYes (opt-in)Grasp-and-release tasks
jerk_rms, psd_lf_hfactionlower / higher is smootherNeverReview only
idle_lead_s, longest_pause_sactionlower is betterYesHuman teleop
idle_trail_s, end_motion_ratioactionlower is betterYes, half weightHuman teleop
length_zmetadataabs(z), lower is betterYesMixed-quality sources
idle_frac, path_length_zactioncontextNeverReview only
track_residaction + statelower is betterYesLeader-follower rigs
accel_spike_frac, joint_limit_fracstatelower is betterYesContact-heavy tasks
track_lag_msaction + statefar from median is flaggedNeverComparing rigs and sessions
gripper_flips_per_s, missed_grasp_fracgripperlower is betterYesGrasp tasks
recovery_countgripperneutralNeverUnderstanding recovery behavior
action_divergencestate + actionlower is betterYes20+ episodes per task
stats_leveragemeta/stats.jsonlower is betterNeverBefore training
blur, exposure_err, clipped_fracvideo framessee aboveYes (camera group)Multi-session data
frozen_fracvideo + statelower is betterYes (camera group)Always, if you can afford to decode
video_action_lag_msvideo + actionnear 0 is in syncYes (camera group)Multi-rig data
task_missing, task_generictask stringflagOwn verdictFine-tuning a VLA
iforest_score, novelty_knnmetric vectorcontextNeverSecond-look queue
Every curation metric in this guide: what it reads, which direction is worse, whether it enters the score, and when to use it.

Beyond heuristics: the model-based metrics

Everything above runs without a model. That’s useful: nothing to download, nothing to host, and cheap, repeatable results. It also means certain questions stay unanswered. A heuristic can’t tell you that an instruction describes the wrong video, that two episodes are near-duplicates in what they show, or whether the robot actually completed the task.
The research literature has methods for each. They cost more, and their evidence is a mix of strong and thin.
One thing to be clear about: the plugin this post demonstrates computes none of these. It’s deliberately heuristic-only. FiftyOne provides the infrastructure for the model-based tier. It has embeddings and similarity through the FiftyOne Brain, a model zoo (including remote zoo models) for running vision-language models (VLMs), and the same panel and tagging workflow to review whatever those produce. You’d build these as additional operators on the same pattern.
1. Instruction-video agreement (VLM). Sample frames, ask a vision-language model whether the video shows the instruction, and store the answer, a confidence score, and a proposed rewrite. SmolVLA’s data pipeline did a version of this, using Qwen2.5-VL-3B to rewrite task strings from sample frames plus the original label, with a prompt asking for under 30 characters starting with an action verb (“Pick,” “Place,” “Open”). Watch out: VLM judges are noisy, and OpenGVL’s benchmark found open VLMs lag proprietary ones. Keep the result as its own verdict, never folded into a score.
2. Embedding-based duplicates and novelty. Embed sampled frames and the action trajectory, then find near-duplicates by cosine similarity. SCIZOR does this at the transition level, using a joint state-action embedding, and reports that its two thresholds (εs = 0.58, εd = 0.99) transferred across RoboMimic, Open X-Embodiment, and real data. Watch out: near-duplicates in a table-top task might be exactly the repetition that robust policies need. Duplication needs a human decision about how much is too much.
3. Task progress. Ask a model to order shuffled frames by progress, then measure how well that order matches the true frame order (Value-Order Correlation, from GVL). OpenGVL applied this to LeRobot Hub data and found task-definition problems, labeling ambiguity, and failed or out-of-distribution episodes. TOPReward replaces generated numbers with token probabilities and reports VOC of 0.874 with Qwen3-VL-8B against 0.218 for GVL with the same model on a filtered Open X-Embodiment subset. Watch out: these are model-dependent, and the numbers come from the authors’ own benchmarks.
4. Information-based scoring. DemInf (Hejna et al., RSS 2025) estimates each demonstration’s contribution to state-action mutual information with k-nearest neighbor (kNN) estimators over learned embeddings. Its abstract reports a 5-10% improvement on RoboMimic and better performance on real ALOHA and Franka setups. The authors note it works best with relative actions.
5. Policy-in-the-loop. CUPID uses influence functions to estimate each demonstration’s effect on expected return from evaluation rollouts, and reports that under 33% of curated data yields state-of-the-art RoboMimic diffusion policies. Demo-SCORE trains a classifier on successful versus failed policy rollouts and filters demonstrations that resemble the failures, and its abstract reports over 15-35% higher absolute success. These need a policy and rollouts, so they’re the highest-fidelity and the most expensive tier.
The CUPID finding to remember. In CUPID’s results, DemInf curated the “highest overall quality,” yet CUPID-curated policies matched or beat it. The authors write that “human perception of demonstration quality does not necessarily correspond to data that maximizes downstream policy success.” That’s a direct warning about every heuristic score in this guide, including the ones that look best.

Dataset-level metrics

Some questions aren’t about any one episode.
  • Task coverage. Episodes per canonical task. A dataset with 1,900 episodes of one task and 12 of another isn’t balanced, no matter what the total says.
  • Source balance. Episodes per source repo when you’ve merged several.
  • Diversity. How spread out trajectories are within a task. Belkhale et al. also found “state diversity is not always beneficial,” so more spread isn’t automatically better.
  • Mixture weights. If you’re combining datasets, the weights matter as much as the filtering. Re-Mix (Hejna et al., CoRL 2024) optimizes domain weights and reports beating uniform weights by 38%. Octo and OpenVLA both curated mixtures from Open X-Embodiment by hand. Octo dropped datasets without images or delta end-effector control, plus ones “too repetitive, low image resolution, or excessively niche.” OpenVLA kept single-arm, third-person-camera, end-effector datasets.
The panel reports coverage as a count on the Integrity & Coverage tab. It’s a count, not a pick, since the right balance depends on what you’re training for.

From metrics to a ranking

Dozens of metrics don’t help you decide where to look. One ranking does. Getting there takes two decisions: compared to what, and how to combine them.

Compared to what: normalization

Task strings are canonicalized, then episodes are grouped by task (an episode with several tasks is grouped by its sorted set). Groups with at least 20 episodes are scored against themselves. A smaller group is pooled across views, and a view with fewer than 20 episodes overall is pooled and marked low-confidence in the panel.
The scoring uses robust z-scores with a floored scale. Zero-inflated metrics (such as idle_lead_s and frame_gaps, where most episodes score exactly zero) use percentile ranks or a tail-based scale, because a standard deviation computed over mostly zeros is meaningless.
Why per task. Without it, you rank tasks against each other. When we ran an earlier version of this scoring on a multi-task robot dataset, two tasks topped the ranking as a group. That’s the cross-task effect: a batch that mixes tasks produces misleading outliers, because “normal” differs by task. Score episodes of the same task together.
Figure 9. A slower, smaller task (B) pooled with a larger one (A). Pooled, the warn line sits near A's distribution and flags 26 of B's 30 episodes for being task B. Per task, each group is flagged against its own distribution: 2 of 120 in A and 1 of 30 in B. Synthetic data, robust z with a median and MAD scale.

Combined how: groups, and the max

Correlated metrics share a group, and a group counts as one vote:
GroupScored members
motionsparc, ldlj, sparc_phase
timeidle_lead_s, idle_trail_s, longest_pause_s, end_motion_ratio, length_z
trackingtrack_resid, accel_spike_frac, joint_limit_frac
grippergripper_flips_per_s, missed_grasp_frac
consistencyaction_divergence
camera (opt-in)blur, exposure_err, clipped_frac, frozen_frac, video_action_lag_ms
The scored members of each metric group. Each group counts as one vote in the episode score.
Inside a group, the value is the maximum of the weighted, oriented member z-scores. Across groups, the episode’s score is the maximum group value, with the weighted mean as a tie-breaker. Weights are 1.0 except ldlj, idle_trail_s, and end_motion_ratio at 0.5.
The effect is that five fine ones can’t average away one severe problem in one dimension. An episode with a frozen camera and perfect motion still ranks near the top. Each episode also carries n_flags (how many groups are flagged) and driver (which group is behind its score), so you know why it ranks where it does.
Figure 10. An example episode with made-up z-scores. Every group is fine except camera, where frozen_frac is at 4.3. The score is the maximum group value, n_flags counts groups at or above the warn line (z = 2), and driver names the group behind the score.

The workflow, step by step

A metric earns its place by changing what you do next. Here’s the loop.

Step 1: Import the dataset

One sample is one episode. Set VFF_MULTIMODAL=1 before importing FiftyOne in every process that touches a LeRobot dataset, since it enables the multimodal episode viewer. Install the plugin with the plugin manager or link a checkout into your plugins folder.
import fiftyone as fo

dataset = fo.Dataset.from_dir(
    dataset_dir="/data/lerobot/my_dataset",
    dataset_type=fo.types.LeRobotDataset,
)

# Merge another source if you want to curate across datasets
dataset.add_dir(
    dataset_dir="/data/lerobot/other_dataset",
    dataset_type=fo.types.LeRobotDataset,
)

session = fo.launch_app(dataset)
See the App guide for sidebar filters and Spaces, which is where the curation panel opens.

Step 2: Compute quality

Open the operator browser and run LeRobot curation: compute quality. It’s delegated by default, so it runs in the background. The form has four tabs: Data (what you tell the tool), Metrics (a checkbox per metric, grouped by family), Camera, and Normalization (the smallest task group scored on its own, default 20).
Your picks are remembered, so re-running doesn’t start from scratch.

Step 3: Read the Overview

Open a panel and choose LeRobot Curation. The Overview tab has the score histogram, verdict counts for Integrity and Language, episodes per task, the outlier scatter, and the worst-first ranking. The panel follows the view: filter the grid and it re-ranks.

Step 4: Drill into the family that drives the problem

Use the driver column to pick a tab. Motion & Action shows one histogram per metric. Vision shows the five camera metrics with camera chips. Click any bar and the samples grid filters to those episodes.

Step 5: Open the inspector

Click any row. The inspector shows every metric with its per-arm or per-camera breakdown, joint traces (action solid, state dashed), the speed profile with the idle threshold, the gripper timeline, and a few frames from the cameras, picked around the flagged spans.
Flagged spans (idle stretches, the longest pause, the roughest smoothness window, and acceleration spikes) are also written as temporal tags on the episode’s timeline once their metric reaches warn so that you can scrub straight to them. Regrasp recoveries are tagged too, as information. A re-run replaces the plugin’s tags and leaves any tag you drew yourself alone.

Step 6: Tag, don’t delete

The footer has three tag buttons: review, exclude-candidate, and relabel. They tag the selection, or everything in view if nothing is selected. Nothing is deleted or hidden. These are ordinary FiftyOne sample tags, so you can filter on them with view stages, save a view, or act on them from code.
Marking an episode exclude-candidate is always a separate human action. The scores never do it for you.

Step 7: Export the kept subset

keep = dataset.match_tags("exclude-candidate", bool=False)

keep.export(
    export_dir="/data/lerobot/my_dataset_curated",
    dataset_type=fo.types.LeRobotDataset,
    export_media=True,
)
The native exporter (API reference) writes a self-contained LeRobot v3 dataset. export_media=True is required, episodes are renumbered contiguously, and task and frame indexes are rebuilt. From there, you can push it to the Hub as you would any LeRobot dataset.
If you’d rather edit in place, send the excluded episode_index values to lerobot-edit-dataset with the delete_episodes operation. Task string rewrites are a separate matter: as of the sources we reviewed, lerobot-edit-dataset can’t yet edit task strings, so accepted rewrites need to be patched into meta/tasks.parquet and the per-frame task_index after export.

Keep it honest

Scores are triage, not verdicts. Smoothness and timing say nothing about whether the demonstration did the right thing. The plugin’s own documentation puts it this way: use it to decide where to look first, and don’t use it as an automatic accept/reject gate.
The evidence is genuinely mixed.
  • Smoothness filtering helped in RINSE, but not in behavior cloning with proprioceptive signals on episodes already filtered for success.
  • The curation-metrics audit found 5 of 7 metrics exploited episode length, and detection accuracy didn’t predict policy success. That’s one preprint on 80 scripted demonstrations. It’s also exactly why outlier metrics aren’t scored and why the length check exists.
  • CUPID found that data humans would call high quality can underperform data selected by influence on the policy.
  • The gripper and tracking metrics have no published downstream validation.
The thresholds are starting points. Defaults such as the 2% joint-limit margin, the warn line at z ≥ 2, the 20-episode group size, and the camera thresholds were calibrated on a 102-episode development set. They aren’t yet adjustable in the form. Check them on your own data.
How the metrics are checked. Each scored metric has a corruption it must catch: added hesitation, an inserted idle run, a duplicated frame, a blanked task, a stalled joint, and a blurred or frozen video. The harness applies the corruption to a real episode and asserts that the corrupted copy ranks worse than the original. It also re-scores every episode truncated to a common length, to find metrics that only measure duration. On 102 episodes, it takes about 5 minutes with camera metrics included. That tells you a metric detects what it claims to detect. It can’t tell you the detection predicts policy quality.
So validate. Before you drop data, train on the curated subset and the full set, then compare. A curation run you haven’t validated with a training run is a hypothesis.

Where FiftyOne fits

Rerun and Foxglove are excellent at what they’re built for: looking closely at a recording. If you want a synchronized view of cameras, joint traces, and 3D for one episode, they’re good tools.
Curation asks a different question, which is about a collection. Which of these 10,000 episodes should I look at first? Of those, which ones have a camera problem? Show me only the ones from this session. Tag these 40. Give me everything else as a dataset I can train on.
That needs a dataset model rather than a viewer: episodes as queryable samples, fields you can filter on, charts that filter the grid when you click them, tags that persist, saved views, and an export that round-trips to the training format. FiftyOne is built around that model, and that’s why the workflow above works the way it does.
The part that matters most if you build your own is how little of this is fixed. FiftyOne’s plugin framework is Python operators plus React panels, and you can mix the two in hybrid panels. The panel in this guide is a React app backed by operators. One data operator returns the whole payload (rows, histogram bins, thresholds, and verdict counts) and re-fires whenever the view changes. A second operator returns a single episode’s arrays and frames on demand for the inspector. The five tabs are handwritten React. Adding a metric is one Python function and one dictionary entry, and charting it means copying a documented chart template.
If your rig has a metric nobody else computes, such as force-torque spikes, a custom contact signal, or a per-operator score, you add it, and it gets the same filter, tag, inspector, and export workflow as everything else.
Nothing in this post requires this exact plugin. It’s one implementation of the pattern.

Try it yourself

Frequently asked questions

Headshot of Harpreet Sahota
Harpreet Sahota
Hacker-in-Residence
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.
See all articles by Harpreet Sahota

Talk to an AI expert

Loading related posts...