Your Robot Data Needs Custom Tools. Build Them as FiftyOne Plugins
Oct 6, 2026
•
13 min read
Author
Harpreet Sahota
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.
How the FiftyOne plugin framework turns your MCAP and LeRobot curation checks into buttons, panels, and background jobs that run on every episode, and why Rerun, Foxglove, and Encord can't host that kind of tool the same way.
You have 400 teleoperation episodes and a policy to train by Friday.
Some of those episodes are bad. The operator hesitated for eight seconds before the grasp. The wrist camera froze halfway through. The gripper opened and closed three times because the first grasp slipped. The task string says "pick up" and nothing else.
A policy will learn from all of it.
You could open each episode in a log viewer and scrub through it. At two minutes per episode, that's more than 13 hours, and you'll stop paying attention long before the end. What you want is a ranked list, worst first, with every flag one click away from the moment it happened.
No viewer ships that list, because the checks depend on your robot. A dual-arm rig with multi-joint hands needs different smoothness metrics than a wheeled rover with an inertial measurement unit (IMU) and GPS. The tool that curates your data has to be one you can extend in Python, on your schema, across the whole dataset.
A FiftyOne plugin is a folder with two files, fiftyone.yml and __init__.py. FiftyOne discovers it at runtime, so you never fork FiftyOne or wait for a release to add a button, form, panel, or background job.
FiftyOne reads .mcap, .bag, and .rrd files as media_type="multimodal" episodes and imports LeRobot v3 with fo.types.LeRobotDataset, so a plugin can score every episode in a dataset with the same code.
FiftyOne plugin code runs on the FiftyOne server in Python so that it can use NumPy, SciPy, PyTorch, a GPU, or a secret API key. Long jobs run in the background through delegated execution.
Rerun extends its viewer in Rust only, with interfaces its docs say break every release. Foxglove extensions are TypeScript running in the browser against one recording. Encord custom agents are HTTPS endpoints tied to labeling workflow stages. None of the three gives you server-side Python tools with custom UI that operate on a whole dataset.
Two open-source plugins show the pattern for physical AI: lerobot_data_curation ranks LeRobot episodes worst-first, and demo_quality_scorer does the same for MCAP recordings.
Why physical AI data needs custom tooling
Image datasets are forgiving. Every sample is a picture, every label is a box or a mask, and the same quality checks work on most of them.
Robot data has no such standard shape. One MCAP file from an autonomous vehicle might hold six cameras, two LIDARs, an IMU, GPS fixes, and a /tf tree. One LeRobot episode from a tabletop arm holds two camera streams, a 14-dimensional observation.state, a matching action array, and a task string. An egocentric dataset adds audio and hand poses.
The questions you need answered are specific to that shape:
Is the motion in this episode smooth, or full of corrections?
Did any sensor drop frames, drift its clock, or clip its values?
Does the action track the state, or is there lag between video and action?
Is the gripper behavior consistent with the other episodes of the same task?
Is the language instruction usable for fine-tuning a vision-language-action (VLA) model?
A generic viewer can't answer those. It can only show you the data and wait for you to notice.
FiftyOne gives you the dataset layer: episodes as samples, fields you can filter and sort on, views, tags, and an App with a multimodal viewer that plays MCAP and LeRobot streams in lockstep. Plugins give you the part that's specific to your robot.
Diagram of FiftyOne at the center of an ML stack, with plugin operators and panels connecting it to labeling tools, vector search backends, model inference, and automation.
What is a FiftyOne plugin?
A plugin is a folder FiftyOne finds. If the folder contains a fiftyone.yml, FiftyOne loads it at runtime and adds whatever it declares to the App, the SDK, and the CLI.
Plugins add two kinds of things:
Operators are actions. You trigger one, it runs, it finishes. "Compute episode quality" is an operator.
Panels are UI that stays open next to your data and reacts to what you do. A worst-first ranking that re-sorts when you filter the grid is a panel.
Under the hood, a panel is a special kind of operator, so everything below about operators applies to panels too.
Notebook script
FiftyOne plugin operator
Who can run it
You, in your kernel
Anyone with the App open, no code
Inputs
Variables you edit by hand
A typed form, validated before anything runs
What it acts on
The dataset you loaded
The view or selection the user has open
Long jobs
Your kernel is busy until it finishes
Delegated to a worker, so the App stays free
Reuse
Copy and paste
One URI: buttons, panels, pipelines, notebooks, agents
What changes when you promote a notebook script to a FiftyOne plugin operator.
Good to know. Why not keep it in a notebook? Start in a notebook. Promote the code to an operator when someone else needs to run it, or when you catch yourself running it a second time. The figure below shows what changes.
Comparison of a notebook script and a FiftyOne plugin operator across who can run it, inputs, what it acts on, long jobs, and reuse.
The row that matters most for robot data is "what it acts on". A script scores whatever dataset you loaded. An operator scores the view the user has open right now, so the same code works on one task, one robot, one week of collection, or the episodes someone tagged for review.
The anatomy of a FiftyOne plugin
Here's the folder for a plugin with a Python backend and a custom React panel:
FiftyOne plugin folder with required fiftyone.yml and _init_.py files and optional requirements.txt and JavaScript files, showing what loads on the Python server and what loads in the App.
fiftyone.yml is the manifest. It holds the plugin name in @org/name form, the operators and panels it exposes, and any secrets it needs.
__init__.py holds your classes and a register(p) function.
requirements.txt lists pip dependencies. For MCAP work, that's usually mcap, mcap-protobuf-support and friends.
Everything below the dashed line is optional JavaScript for custom UI.
A Python-only plugin is two files. Teal in every figure in this post means the FiftyOne server, which runs Python. Orange means the App in your browser.
Keep it honest. Every operator name listed in fiftyone.yml must match the name in that operator's config. If you register a class but forget to list it, the operator doesn't show up, and you get no error. It's the most common "where did my operator go" bug.
A complete operator in 25 lines
This operator tags samples whose detections fall below a confidence threshold:
TagLowConf operator code color-coded to the button, input form, and notification it creates in the FiftyOne App.
Four methods, four jobs:
config names the operator. name is its ID, and label becomes the form's heading.
resolve_placement puts a "Tag low-conf" button in the grid's action bar.
resolve_input declares a dropdown and a number. The App builds the form, with defaults and validation, from those two lines.
execute does the work on the target view and sends a notification.
You wrote no UI code. You described the inputs and the App built the interface. Swap the body of execute for spectral arc length (SPARC) smoothness over an /odom channel, and you have the skeleton of an MCAP quality scorer.
How an operator call works
Every operator call is two round trips between the App and the FiftyOne server:
Sequence diagram of a FiftyOne operator call, with one round trip that resolves inputs into a form and a second that executes and returns UI operations.
Resolve inputs. You click the button. The App sends the current context (dataset, view, selection) to the server, which calls your resolve_input and returns a schema. The App turns it into a form. If the operator is dynamic, this loop re-runs on every input change, which is how demo_quality_scorer shows a signal picker only after you choose a telemetry channel.
Execute. You hit Execute. The server validates the parameters against your schema, then calls execute. You return a result plus any UI operations, and the App applies them.
Your code runs only on the server. The browser draws what the server tells it to. That's why an operator can decode protobuf MCAP messages with mcap-protobuf-support, run SciPy's Welch power spectral density (PSD) on an IMU stream, or call a model on a GPU, without shipping any of that to the browser.
When something breaks, ask which arrow failed. A wrong form means resolve_input. A fine form that does nothing means execute. An execute that ran but didn't change the App means the UI operations you returned.
Long jobs run in the background with delegated execution
Scoring 400 episodes is not a click-and-wait job. Decoding video alone costs lerobot_data_curation about 0.4 seconds per camera per episode.
Flowchart comparing immediate execution in the FiftyOne server with delegated execution, where jobs queue in MongoDB and an orchestrator runs them.
Set allow_delegated_execution=True in the operator's config, and the user can choose. Immediate execution runs inside the server while the App waits, with a live progress bar. Delegated execution writes the job to a queue in MongoDB, and a separate worker runs the same execute method. You start a worker with fiftyone delegated launch and track jobs with fiftyone delegated list.
Both physical AI plugins below default to delegated execution, because a real corpus takes minutes.
Keep it honest. If no worker is running, delegated jobs sit in the queue forever. And a delegated job has no browser attached, so write results to the dataset rather than trying to change the user's view.
Panels: UI that stays open
A curation tool needs more than a form. It needs a place to see the ranking, click into an episode, and tag it. That's a panel.
FiftyOne panel lifecycle from on_load to render to event handlers, next to four places to keep state: panel state, panel data, execution store, and dataset fields.
A panel's on_load runs when it opens, render returns a layout, and events call your Python handlers with a fresh context. The event that matters most for curation is on_change_view: when the user filters the grid, your panel re-computes for the new view.
Here's a complete panel that shows class counts and follows the grid:
Small values go in panel state, which your Python handlers can read. Large payloads like chart data go through set_data, which is write-only from Python. Settings you want to survive a reload go in the execution store, ctx.store. Results that belong to the data go in dataset fields, where you can filter and sort on them like any other field.
When the built-in widgets run out, you have two options. A hybrid panel keeps state and compute in Python and renders a React component you write. A JS panel puts React in charge of the UI and calls your Python operators whenever it needs the server. Both plugins below are JS panels backed by Python operators, which suits dense, interactive dashboards. If you build a React UI, VOODO, the component library the FiftyOne App itself uses, makes it look native.
Comparison of Python, hybrid, and JS FiftyOne panels showing what runs in the browser, what runs on the server, and which need a Node build.
Operators compose
Every operator has a URI, like @voxel51/brain/compute_similarity, and a typed input schema. Anything that can call an operator can call yours.
One FiftyOne operator called from eight places: an App button, the operator browser, another operator, a pipeline stage, a notebook, AI agents through the MCP server, a background worker, and teammates.
That includes the FiftyOne Model Context Protocol (MCP) server, which exposes list_operators, get_operator_schema, and execute_operator to AI agents. An agent can find your quality scorer, read its inputs, run it on the episodes tagged review, and report back which ones failed sensor health. The schema validates the agent's JSON the same way it validates a person filling in the form.
For a physical AI team, that means the curation logic you write once runs from a button during review, from a notebook in your training pipeline, from a background worker overnight, and from an agent.
How FiftyOne plugins compare to other extension points
Rerun and Foxglove are good log viewers. Encord is a labeling platform with curation features. All three have extension points. The difference is what those extension points can reach.
How FiftyOne, Rerun, Foxglove, and Encord let you extend the tool: language, where code runs, dataset scope, and custom UI.
Rerun. Python can arrange Rerun's built-in views through blueprints, but adding a new view, visualizer, or panel means writing Rust against re_viewer and shipping your own viewer binary. Rerun's own docs warn: "The interfaces for extending the Viewer are not yet stable. Expect code implementing custom extensions to break with every release of Rerun." That's a reasonable trade-off for a young viewer. It's a hard foundation for a team curation tool you want to maintain for a year.
Foxglove. Foxglove has the most complete extension system of the three. You can write custom panels in React, convert custom message schemas so built-in panels can display them, and add data loaders for new file formats. All of it is TypeScript running in the browser or desktop app, scoped to the data source you have open. There's no server-side Python step that scores 400 recordings, writes the results somewhere you can sort on, and hands you back a ranked list. Installing local extensions also requires a developer seat.
Encord. Encord's custom agents let you run your own Python on a task: register an HTTPS endpoint, receive the project, data, and frame as JSON, write labels back through the SDK. Task agents move work between workflow stages. That's useful for pre-labeling and QA routing. An agent can't add a panel, a chart, or a new kind of form to Encord's interface, and it acts on tasks in a labeling workflow rather than on an arbitrary filtered view of your data.
FiftyOne's plugin framework is the only one of the four where a single Python file can add a form, run a job across every episode in the background, render a custom panel with your results, and expose all of it to agents, without leaving the language your robotics stack is written in.
Keep it honest. For live debugging of a single recording, scrubbing a running robot's topics in real time, Foxglove and Rerun are excellent, and FiftyOne doesn't try to replace them. FiftyOne reads .rrd files and MCAP files produced by Foxglove tooling, so the common setup is to debug individual logs in those viewers and curate the dataset in FiftyOne.
Two FiftyOne plugins for MCAP and LeRobot curation
Both plugins below were built using only the pieces covered above. Both follow the same rule, stated in their READMEs: the scores order your review queue, and a person makes the call. A perfectly smooth demonstration can still show the wrong task.
It scores each episode on motion smoothness, time efficiency, action-state tracking, gripper behavior, consistency, integrity, and language, and optionally on camera quality: blur, exposure, clipping, frozen feeds, and video-action lag. It has no model-based metrics, so there's no model to download.
How it maps to the framework:
A typed form with tabs. The compute operator's form has Data, Metrics, Camera, and Normalization tabs. You pick the observation.state and action arrays, and assign arm and gripper joints by name. Nothing is guessed from the data, and your picks are remembered for the next run.
Delegated by default. The operator sets allow_delegated_execution=True and default_choice_to_delegated=True, so scoring runs in the background. Camera decoding is the slow part.
A JS panel backed by Python operators. The LeRobot Curation panel is a React component with Overview, Motion & Action, Integrity & Coverage, Vision and Language tabs. Click a histogram bar to filter the grid to those episodes. Click a row to open an inspector with joint traces, the speed profile, the gripper timeline, and frames picked around the flagged spans.
It follows the view. The panel refreshes when the view changes, so filtering the grid re-ranks it.
Tags on the timeline. Flagged spans such as idle stretches, the longest pause, and acceleration spikes are written as temporal tags on each episode so that you can scrub straight to them in the multimodal viewer. Footer buttons apply review, exclude-candidate, and relabel sample tags.
import fiftyone as fo
dataset = fo.Dataset.from_dir(
dataset_dir="/data/lerobot/my_dataset",
dataset_type=fo.types.LeRobotDataset,
)
session = fo.launch_app(dataset)
# then run "LeRobot curation: compute quality" from the operator browser
Keep it honest. Scores are relative to the batch you scored. Per-task normalization needs about 20 episodes per task, and the default thresholds were calibrated on a 102-episode development set. Check them on your own data. Cloud-hosted LeRobot sources aren't supported yet.
It computes three metric families and ranks episodes worst-first:
Motion smoothness: SPARC, log dimensionless jerk (LDLJ), jerk RMS, and a low-to-high frequency power ratio, computed over windows of a speed profile you choose from position or velocity signals such as /odom -> twist_linear.
Sensor health: dropout, rate stability, clock drift, clipping, and cross-channel desync, all from message timestamps and raw values, so it works on any channel the plugin can decode.
Outliers: models that flag episodes whose motion and health features don't look like the rest of the batch.
How it maps to the framework:
A dynamic form. The form inspects the first sample's channels. If no telemetry channel carries a numeric signal, the motion family switches itself off and says why. If the server lacks a decoder, the form names the package to install.
Immediate or delegated, your choice. Run immediately with an in-app progress bar, or schedule it in the background (default).
A JS panel. The Episode Quality panel is a React component that has Motion, Health, and Outliers tabs. Click a histogram bar to filter the grid. Click a table row to open that episode in the multimodal viewer, with a toast naming the timecode of its worst flagged interval.
Results as fields. Scores land in quality.* fields such as quality.sparc and quality.ldlj, so you can sort and filter on them outside the panel too.
Keep it honest. Structural errors, such as the wrong action at a key moment, are invisible to every metric here. Re-running over a different set of episodes re-fits normalization, so expect scores to move.
Build your own in six steps
Your robot has checks neither plugin covers. Here's the loop for writing them.
Six steps to build a FiftyOne plugin: scaffold, write, register, run, iterate, and ship, with a loop from iterate back to write.
Step 1: Scaffold the plugin
Run fiftyone plugins create or make a folder in your plugins directory with a fiftyone.yml:
One class per operator or panel, like the examples above. Keep the scoring logic in a plain Python function so a notebook, a test, and the operator can all call it.
Step 3: Register it
Call p.register(ScoreGrasps) inside register(p), and make sure the name matches fiftyone.yml.
fiftyone app debug prints server logs and tracebacks to your terminal. VFF_MULTIMODAL=1 turns on the multimodal viewer in the App; set it before fiftyone is imported in every process that touches a multimodal dataset.
Step 5: Iterate
Edit, refresh the App, repeat. Python changes are picked up on the next request without a server restart. If you have a React panel, keep npm run dev rebuilding it.
Good to know. How do I test an operator without the App? Operators are callable from Python with foo.execute_operator, so you can test them headless in pytest.
Import it with fo.Dataset.from_dir(dataset_dir=..., dataset_type=fo.types.LeRobotDataset), where each sample is one episode. Then install lerobot_data_curation and run its compute operator to rank episodes worst-first on motion, gripper, camera, and language metrics. Review the top of the list and tag episodes as exclude-candidate or relabel.
Yes. FiftyOne treats .mcap, .bag, and .rrd files as media_type="multimodal" samples, one episode per file, and plays their streams in a synchronized multimodal viewer. Set VFF_MULTIMODAL=1 before importing fiftyone to turn the viewer on.
Score every episode on motion smoothness (SPARC, LDLJ, jerk) and sensor health (dropout, rate stability, clock drift), then sort worst-first. demo_quality_scorer does this for MCAP datasets. Treat the scores as a review queue, because a smooth demonstration can still show the wrong task.
Use Foxglove to debug a single recording, especially live robot data. Use FiftyOne to curate a dataset of recordings: filter, sort, score, and tag thousands of episodes, and extend that workflow with Python plugins. Many teams use both, and FiftyOne reads MCAP files produced by Foxglove tooling.
Rerun is a fast viewer for multimodal logs, and FiftyOne reads its .rrd files. Extending Rerun's viewer requires Rust and a custom build, and its interfaces change between releases. FiftyOne plugins are Python, load at runtime, and work across a whole dataset.
Encord's extension point is custom agents: HTTPS endpoints or Python task runners that act on tasks in a labeling workflow and write labels back through the SDK. They don't add panels or new UI to Encord. In FiftyOne, a panel is a Python class a plugin registers.
No. A Python-only plugin is fiftyone.yml plus __init__.py, and the App builds forms and panels from Python declarations. Add React only when you need custom interactivity, such as the dashboards in lerobot_data_curation and demo_quality_scorer.
It runs on the FiftyOne server, in Python, or on a background worker for delegated operations. It never runs in the browser, so it can use any Python library, a GPU, or secrets without exposing them.
Harpreet Sahota
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.