Your Robot Data Needs Custom Tools. Build Them as FiftyOne Plugins

Oct 6, 2026
•
13 min read
How the FiftyOne plugin framework turns your MCAP and LeRobot curation checks into buttons, panels, and background jobs that run on every episode, and why Rerun, Foxglove, and Encord can't host that kind of tool the same way.
You have 400 teleoperation episodes and a policy to train by Friday.
Some of those episodes are bad. The operator hesitated for eight seconds before the grasp. The wrist camera froze halfway through. The gripper opened and closed three times because the first grasp slipped. The task string says "pick up" and nothing else.
A policy will learn from all of it.
You could open each episode in a log viewer and scrub through it. At two minutes per episode, that's more than 13 hours, and you'll stop paying attention long before the end. What you want is a ranked list, worst first, with every flag one click away from the moment it happened.
No viewer ships that list, because the checks depend on your robot. A dual-arm rig with multi-joint hands needs different smoothness metrics than a wheeled rover with an inertial measurement unit (IMU) and GPS. The tool that curates your data has to be one you can extend in Python, on your schema, across the whole dataset.
That's what the FiftyOne plugin framework is for.

Key takeaways

  • A FiftyOne plugin is a folder with two files, fiftyone.yml and __init__.py. FiftyOne discovers it at runtime, so you never fork FiftyOne or wait for a release to add a button, form, panel, or background job.
  • FiftyOne reads .mcap, .bag, and .rrd files as media_type="multimodal" episodes and imports LeRobot v3 with fo.types.LeRobotDataset, so a plugin can score every episode in a dataset with the same code.
  • FiftyOne plugin code runs on the FiftyOne server in Python so that it can use NumPy, SciPy, PyTorch, a GPU, or a secret API key. Long jobs run in the background through delegated execution.
  • Rerun extends its viewer in Rust only, with interfaces its docs say break every release. Foxglove extensions are TypeScript running in the browser against one recording. Encord custom agents are HTTPS endpoints tied to labeling workflow stages. None of the three gives you server-side Python tools with custom UI that operate on a whole dataset.
  • Two open-source plugins show the pattern for physical AI: lerobot_data_curation ranks LeRobot episodes worst-first, and demo_quality_scorer does the same for MCAP recordings.

Why physical AI data needs custom tooling

Image datasets are forgiving. Every sample is a picture, every label is a box or a mask, and the same quality checks work on most of them.
Robot data has no such standard shape. One MCAP file from an autonomous vehicle might hold six cameras, two LIDARs, an IMU, GPS fixes, and a /tf tree. One LeRobot episode from a tabletop arm holds two camera streams, a 14-dimensional observation.state, a matching action array, and a task string. An egocentric dataset adds audio and hand poses.
The questions you need answered are specific to that shape:
  • Is the motion in this episode smooth, or full of corrections?
  • Did any sensor drop frames, drift its clock, or clip its values?
  • Does the action track the state, or is there lag between video and action?
  • Is the gripper behavior consistent with the other episodes of the same task?
  • Is the language instruction usable for fine-tuning a vision-language-action (VLA) model?
A generic viewer can't answer those. It can only show you the data and wait for you to notice.
FiftyOne gives you the dataset layer: episodes as samples, fields you can filter and sort on, views, tags, and an App with a multimodal viewer that plays MCAP and LeRobot streams in lockstep. Plugins give you the part that's specific to your robot.
Diagram of FiftyOne at the center of an ML stack, with plugin operators and panels connecting it to labeling tools, vector search backends, model inference, and automation.

What is a FiftyOne plugin?

A plugin is a folder FiftyOne finds. If the folder contains a fiftyone.yml, FiftyOne loads it at runtime and adds whatever it declares to the App, the SDK, and the CLI.
Plugins add two kinds of things:
  • Operators are actions. You trigger one, it runs, it finishes. "Compute episode quality" is an operator.
  • Panels are UI that stays open next to your data and reacts to what you do. A worst-first ranking that re-sorts when you filter the grid is a panel.
Under the hood, a panel is a special kind of operator, so everything below about operators applies to panels too.
Notebook scriptFiftyOne plugin operator
Who can run itYou, in your kernelAnyone with the App open, no code
InputsVariables you edit by handA typed form, validated before anything runs
What it acts onThe dataset you loadedThe view or selection the user has open
Long jobsYour kernel is busy until it finishesDelegated to a worker, so the App stays free
ReuseCopy and pasteOne URI: buttons, panels, pipelines, notebooks, agents
What changes when you promote a notebook script to a FiftyOne plugin operator.
Good to know. Why not keep it in a notebook? Start in a notebook. Promote the code to an operator when someone else needs to run it, or when you catch yourself running it a second time. The figure below shows what changes.
Comparison of a notebook script and a FiftyOne plugin operator across who can run it, inputs, what it acts on, long jobs, and reuse.
The row that matters most for robot data is "what it acts on". A script scores whatever dataset you loaded. An operator scores the view the user has open right now, so the same code works on one task, one robot, one week of collection, or the episodes someone tagged for review.

The anatomy of a FiftyOne plugin

Here's the folder for a plugin with a Python backend and a custom React panel:
FiftyOne plugin folder with required fiftyone.yml and _init_.py files and optional requirements.txt and JavaScript files, showing what loads on the Python server and what loads in the App.
  • fiftyone.yml is the manifest. It holds the plugin name in @org/name form, the operators and panels it exposes, and any secrets it needs.
  • __init__.py holds your classes and a register(p) function.
  • requirements.txt lists pip dependencies. For MCAP work, that's usually mcap, mcap-protobuf-support and friends.
  • Everything below the dashed line is optional JavaScript for custom UI.
A Python-only plugin is two files. Teal in every figure in this post means the FiftyOne server, which runs Python. Orange means the App in your browser.
Keep it honest. Every operator name listed in fiftyone.yml must match the name in that operator's config. If you register a class but forget to list it, the operator doesn't show up, and you get no error. It's the most common "where did my operator go" bug.

A complete operator in 25 lines

This operator tags samples whose detections fall below a confidence threshold:
import fiftyone as fo
import fiftyone.operators as foo
import fiftyone.operators.types as types
from fiftyone import ViewField as F


class TagLowConf(foo.Operator):
    @property
    def config(self):
        return foo.OperatorConfig(name="tag_low_conf", label="Tag low-confidence")

    def resolve_placement(self, ctx):
        return types.Placement(
            types.Places.SAMPLES_GRID_SECONDARY_ACTIONS,
            types.Button(label="Tag low-conf"),
        )

    def resolve_input(self, ctx):
        inputs = types.Object()
        fields = list(ctx.dataset.get_field_schema(embedded_doc_type=fo.Detections))
        inputs.enum("field", fields, required=True, label="Label field")
        inputs.float("thresh", default=0.3, label="Confidence threshold")
        return types.Property(inputs)

    def execute(self, ctx):
        field, thresh = ctx.params["field"], ctx.params["thresh"]
        view = ctx.target_view().filter_labels(field, F("confidence") < thresh)
        view.tag_samples("low_conf")
        ctx.ops.notify(f"Tagged {len(view)} samples")


def register(p):
    p.register(TagLowConf)
And here's what the user sees:
TagLowConf operator code color-coded to the button, input form, and notification it creates in the FiftyOne App.
Four methods, four jobs:
  1. config names the operator. name is its ID, and label becomes the form's heading.
  2. resolve_placement puts a "Tag low-conf" button in the grid's action bar.
  3. resolve_input declares a dropdown and a number. The App builds the form, with defaults and validation, from those two lines.
  4. execute does the work on the target view and sends a notification.
You wrote no UI code. You described the inputs and the App built the interface. Swap the body of execute for spectral arc length (SPARC) smoothness over an /odom channel, and you have the skeleton of an MCAP quality scorer.

How an operator call works

Every operator call is two round trips between the App and the FiftyOne server:
Sequence diagram of a FiftyOne operator call, with one round trip that resolves inputs into a form and a second that executes and returns UI operations.
  1. Resolve inputs. You click the button. The App sends the current context (dataset, view, selection) to the server, which calls your resolve_input and returns a schema. The App turns it into a form. If the operator is dynamic, this loop re-runs on every input change, which is how demo_quality_scorer shows a signal picker only after you choose a telemetry channel.
  2. Execute. You hit Execute. The server validates the parameters against your schema, then calls execute. You return a result plus any UI operations, and the App applies them.
Your code runs only on the server. The browser draws what the server tells it to. That's why an operator can decode protobuf MCAP messages with mcap-protobuf-support, run SciPy's Welch power spectral density (PSD) on an IMU stream, or call a model on a GPU, without shipping any of that to the browser.
When something breaks, ask which arrow failed. A wrong form means resolve_input. A fine form that does nothing means execute. An execute that ran but didn't change the App means the UI operations you returned.

Long jobs run in the background with delegated execution

Scoring 400 episodes is not a click-and-wait job. Decoding video alone costs lerobot_data_curation about 0.4 seconds per camera per episode.
Flowchart comparing immediate execution in the FiftyOne server with delegated execution, where jobs queue in MongoDB and an orchestrator runs them.
Set allow_delegated_execution=True in the operator's config, and the user can choose. Immediate execution runs inside the server while the App waits, with a live progress bar. Delegated execution writes the job to a queue in MongoDB, and a separate worker runs the same execute method. You start a worker with fiftyone delegated launch and track jobs with fiftyone delegated list.
Both physical AI plugins below default to delegated execution, because a real corpus takes minutes.
Keep it honest. If no worker is running, delegated jobs sit in the queue forever. And a delegated job has no browser attached, so write results to the dataset rather than trying to change the user's view.

Panels: UI that stays open

A curation tool needs more than a form. It needs a place to see the ranking, click into an episode, and tag it. That's a panel.
FiftyOne panel lifecycle from on_load to render to event handlers, next to four places to keep state: panel state, panel data, execution store, and dataset fields.
A panel's on_load runs when it opens, render returns a layout, and events call your Python handlers with a fresh context. The event that matters most for curation is on_change_view: when the user filters the grid, your panel re-computes for the new view.
Here's a complete panel that shows class counts and follows the grid:
class ClassCounts(foo.Panel):
    @property
    def config(self):
        return foo.PanelConfig(name="class_counts", label="Class counts")

    def on_load(self, ctx):
        self.refresh(ctx)

    def on_change_view(self, ctx):
        self.refresh(ctx)

    def refresh(self, ctx):
        counts = ctx.view.count_values("ground_truth.detections.label")
        ctx.panel.set_state("n", len(counts))
        ctx.panel.set_data("counts", [{"type": "bar", "x": list(counts), "y": list(counts.values())}])

    def render(self, ctx):
        panel = types.Object()
        panel.md(f"**{ctx.panel.get_state('n')} classes**")
        panel.plot("counts")
        return types.Property(panel)
Small values go in panel state, which your Python handlers can read. Large payloads like chart data go through set_data, which is write-only from Python. Settings you want to survive a reload go in the execution store, ctx.store. Results that belong to the data go in dataset fields, where you can filter and sort on them like any other field.
When the built-in widgets run out, you have two options. A hybrid panel keeps state and compute in Python and renders a React component you write. A JS panel puts React in charge of the UI and calls your Python operators whenever it needs the server. Both plugins below are JS panels backed by Python operators, which suits dense, interactive dashboards. If you build a React UI, VOODO, the component library the FiftyOne App itself uses, makes it look native.
Comparison of Python, hybrid, and JS FiftyOne panels showing what runs in the browser, what runs on the server, and which need a Node build.

Operators compose

Every operator has a URI, like @voxel51/brain/compute_similarity, and a typed input schema. Anything that can call an operator can call yours.
One FiftyOne operator called from eight places: an App button, the operator browser, another operator, a pipeline stage, a notebook, AI agents through the MCP server, a background worker, and teammates.
That includes the FiftyOne Model Context Protocol (MCP) server, which exposes list_operators, get_operator_schema, and execute_operator to AI agents. An agent can find your quality scorer, read its inputs, run it on the episodes tagged review, and report back which ones failed sensor health. The schema validates the agent's JSON the same way it validates a person filling in the form.
For a physical AI team, that means the curation logic you write once runs from a button during review, from a notebook in your training pipeline, from a background worker overnight, and from an agent.

How FiftyOne plugins compare to other extension points

Rerun and Foxglove are good log viewers. Encord is a labeling platform with curation features. All three have extension points. The difference is what those extension points can reach.
PlatformHow you extend itLanguageWhere your code runsWorks across a whole dataset?Custom UI in the app?
FiftyOnePlugins: operators and panelsPython, optional ReactFiftyOne server and background workersYes, on any viewYes, panels plus buttons in 11 App placements
RerunCustom views and visualizers Rust onlyYour own build of the viewerNoYes, if you compile your own viewer
FoxgloveExtensions: panels, converters, data loadersTypeScript, Rust compiled to WASMBrowser or desktop appNo, one data source at a timeYes, custom panels
EncordCustom agents Python, behind an HTTPS endpointYour hosted endpointPer task in a workflow stageNo, a trigger in the Label Editor
How FiftyOne, Rerun, Foxglove, and Encord let you extend the tool: language, where code runs, dataset scope, and custom UI.
Rerun. Python can arrange Rerun's built-in views through blueprints, but adding a new view, visualizer, or panel means writing Rust against re_viewer and shipping your own viewer binary. Rerun's own docs warn: "The interfaces for extending the Viewer are not yet stable. Expect code implementing custom extensions to break with every release of Rerun." That's a reasonable trade-off for a young viewer. It's a hard foundation for a team curation tool you want to maintain for a year.
Foxglove. Foxglove has the most complete extension system of the three. You can write custom panels in React, convert custom message schemas so built-in panels can display them, and add data loaders for new file formats. All of it is TypeScript running in the browser or desktop app, scoped to the data source you have open. There's no server-side Python step that scores 400 recordings, writes the results somewhere you can sort on, and hands you back a ranked list. Installing local extensions also requires a developer seat.
Encord. Encord's custom agents let you run your own Python on a task: register an HTTPS endpoint, receive the project, data, and frame as JSON, write labels back through the SDK. Task agents move work between workflow stages. That's useful for pre-labeling and QA routing. An agent can't add a panel, a chart, or a new kind of form to Encord's interface, and it acts on tasks in a labeling workflow rather than on an arbitrary filtered view of your data.
FiftyOne's plugin framework is the only one of the four where a single Python file can add a form, run a job across every episode in the background, render a custom panel with your results, and expose all of it to agents, without leaving the language your robotics stack is written in.
Keep it honest. For live debugging of a single recording, scrubbing a running robot's topics in real time, Foxglove and Rerun are excellent, and FiftyOne doesn't try to replace them. FiftyOne reads .rrd files and MCAP files produced by Foxglove tooling, so the common setup is to debug individual logs in those viewers and curate the dataset in FiftyOne.

Two FiftyOne plugins for MCAP and LeRobot curation

Both plugins below were built using only the pieces covered above. Both follow the same rule, stated in their READMEs: the scores order your review queue, and a person makes the call. A perfectly smooth demonstration can still show the wrong task.

LeRobot Data Curation

lerobot_data_curation ranks the episodes of a LeRobot v3 dataset worst-first, so you know which to look at before training a policy or fine-tuning a VLA.
It scores each episode on motion smoothness, time efficiency, action-state tracking, gripper behavior, consistency, integrity, and language, and optionally on camera quality: blur, exposure, clipping, frozen feeds, and video-action lag. It has no model-based metrics, so there's no model to download.
How it maps to the framework:
  • A typed form with tabs. The compute operator's form has Data, Metrics, Camera, and Normalization tabs. You pick the observation.state and action arrays, and assign arm and gripper joints by name. Nothing is guessed from the data, and your picks are remembered for the next run.
  • Delegated by default. The operator sets allow_delegated_execution=True and default_choice_to_delegated=True, so scoring runs in the background. Camera decoding is the slow part.
  • A JS panel backed by Python operators. The LeRobot Curation panel is a React component with Overview, Motion & Action, Integrity & Coverage, Vision and Language tabs. Click a histogram bar to filter the grid to those episodes. Click a row to open an inspector with joint traces, the speed profile, the gripper timeline, and frames picked around the flagged spans.
  • It follows the view. The panel refreshes when the view changes, so filtering the grid re-ranks it.
  • Tags on the timeline. Flagged spans such as idle stretches, the longest pause, and acceleration spikes are written as temporal tags on each episode so that you can scrub straight to them in the multimodal viewer. Footer buttons apply review, exclude-candidate, and relabel sample tags.
import fiftyone as fo

dataset = fo.Dataset.from_dir(
    dataset_dir="/data/lerobot/my_dataset",
    dataset_type=fo.types.LeRobotDataset,
)
session = fo.launch_app(dataset)
# then run "LeRobot curation: compute quality" from the operator browser
Keep it honest. Scores are relative to the batch you scored. Per-task normalization needs about 20 episodes per task, and the default thresholds were calibrated on a 102-episode development set. Check them on your own data. Cloud-hosted LeRobot sources aren't supported yet.

Demo Quality Scorer

demo_quality_scorer does the same job for multimodal MCAP datasets: teleoperations, autonomous vehicles, drones, rovers, and egocentric capture.
It computes three metric families and ranks episodes worst-first:
  1. Motion smoothness: SPARC, log dimensionless jerk (LDLJ), jerk RMS, and a low-to-high frequency power ratio, computed over windows of a speed profile you choose from position or velocity signals such as /odom -> twist_linear.
  2. Sensor health: dropout, rate stability, clock drift, clipping, and cross-channel desync, all from message timestamps and raw values, so it works on any channel the plugin can decode.
  3. Outliers: models that flag episodes whose motion and health features don't look like the rest of the batch.
How it maps to the framework:
  • A dynamic form. The form inspects the first sample's channels. If no telemetry channel carries a numeric signal, the motion family switches itself off and says why. If the server lacks a decoder, the form names the package to install.
  • Immediate or delegated, your choice. Run immediately with an in-app progress bar, or schedule it in the background (default).
  • A JS panel. The Episode Quality panel is a React component that has Motion, Health, and Outliers tabs. Click a histogram bar to filter the grid. Click a table row to open that episode in the multimodal viewer, with a toast naming the timecode of its worst flagged interval.
  • Results as fields. Scores land in quality.* fields such as quality.sparc and quality.ldlj, so you can sort and filter on them outside the panel too.
pip install "fiftyone[multimodal]>=1.19.0" mcap mcap-protobuf-support \
    mcap-ros1-support mcap-ros2-support numpy scipy scikit-learn
fiftyone plugins download https://github.com/harpreetsahota204/demo_quality_scorer
Keep it honest. Structural errors, such as the wrong action at a key moment, are invisible to every metric here. Re-running over a different set of episodes re-fits normalization, so expect scores to move.

Build your own in six steps

Your robot has checks neither plugin covers. Here's the loop for writing them.
Six steps to build a FiftyOne plugin: scaffold, write, register, run, iterate, and ship, with a loop from iterate back to write.

Step 1: Scaffold the plugin

Run fiftyone plugins create or make a folder in your plugins directory with a fiftyone.yml:
name: "@your-org/gripper_checks"
description: Flags episodes with repeated grasp attempts
operators:
  - score_grasps

Step 2: Write the operator

One class per operator or panel, like the examples above. Keep the scoring logic in a plain Python function so a notebook, a test, and the operator can all call it.

Step 3: Register it

Call p.register(ScoreGrasps) inside register(p), and make sure the name matches fiftyone.yml.

Step 4: Run the App in debug mode

export VFF_MULTIMODAL=1
fiftyone app debug my_lerobot_dataset
fiftyone app debug prints server logs and tracebacks to your terminal. VFF_MULTIMODAL=1 turns on the multimodal viewer in the App; set it before fiftyone is imported in every process that touches a multimodal dataset.

Step 5: Iterate

Edit, refresh the App, repeat. Python changes are picked up on the next request without a server restart. If you have a React panel, keep npm run dev rebuilding it.

Step 6: Ship it

Push to GitHub. Anyone installs it with:
fiftyone plugins download https://github.com/your-org/gripper_checks
fiftyone plugins requirements @your-org/gripper_checks --install
Good to know. How do I test an operator without the App? Operators are callable from Python with foo.execute_operator, so you can test them headless in pytest.

Try it yourself

Frequently Asked Questions

Harpreet Sahota avatar
Harpreet Sahota
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.
See all articles by Harpreet Sahota

Talk to an AI expert

Loading related posts...