Annotation built for synchronized, multi-sensor episodes

FiftyOne supports multimodal data annotation with synchronized 2D and 3D annotation. Customize workflows to bridge annotation, data curation, and model evaluation in one unified platform.
Bounding box annotations labeling each fish for object detection in an underwater coral reef image
Image labeling interface applying classification tags like daytime and zebra crossing to a pedestrian street scene, with an approve-labels review step

Smarter Multimodal Labeling

Go from raw sensor recordings to labeled, training-ready episodes

FiftyOne supports every annotation type your physical AI models need, from temporal tags on time intervals to label tracks that persist across the full episode, with synchronized 2D and 3D context throughout.

Temporal Tags

Instantly mark moments that matter
Tag any interval in the recording with a single click. Capture near-misses, sensor dropouts, interesting behaviors, and edge cases for better model training. Tags carry rich metadata, persist alongside your data, and are fully queryable across your entire dataset.

Label Tracks

Annotations that persist across the full episode
Label tracks follow objects and events across the entire recording instead of frame by frame. See when objects appear, behaviors overlap, or events unfold on a shared timeline, and query those same tracks to find similar patterns across your fleet.

2D Camera Annotations

Precise labels on synchronized camera streams
Annotate 2D regions, detections, and classifications directly on camera image streams. Annotations are synchronized with every other sensor in the recording, so a label on a camera frame is always in context with the corresponding lidar scan and sensor data at the exact timestamp you want.

3D Spatial Annotations

3D labels in a shared world frame
Annotate objects and regions in 3D within the synchronized scene viewer. 3D annotations are placed in a shared world frame alongside point clouds, pose trajectories, and camera frustums, so your labels reflect the full spatial context of the recording.

MCAP Ingestion

Every sensor, in sync. No pre-extraction required.

Most teams lose episode context when they split recordings into frames for annotation. FiftyOne doesn't.
FiftyOne ingests MCAP files natively, camera feeds, LiDAR point clouds, radar tracks, IMU readings, GPS, and more play back in a synchronized, tiled viewer without lossy pre-extraction or frame-splitting pipelines. You see exactly what every sensor captured at the precise millisecond of any event. When you tag an interval or draw an annotation, it's timestamped against the shared playback clock and preserved in the context of the full recording.

FiftyOne Multimodal Annotation

Designed for teams building physical AI models at scale

Physical AI annotation breaks in ways 2D tools weren’t made to handle. FiftyOne is designed around the failure modes of episode-level, multi-sensor data.
Ontologies built for episode-level data
Define clean, thoughtful ontologies for creating high-quality labeled image data with consistencies and minimal errors at scale.
Query your fleet for similar episodes
Surface the edge cases and coverage gaps that reveal where your model is failing. Find every episode where a specific behavior occurred by filtering metadata, temporal event, label, or embedding similarity, all without scanning raw MCAP files.
Close the loop without switching tools
When annotation lives alongside curation and model evaluation, your team can isolate a failure, tag the exact temporal slice, query for similar conditions, and easily route them back into model training.

ML Research

Auto-labeling rivals human performance

The latest paper from our ML researchers, Auto-Labeling Data for Object Detection, benchmarks auto-labeling against human annotation. We reveal how foundation models can deliver labels at near-human accuracy, while reducing annotation costs by up to 100,000×.
The latest paper from our ML researchers, Auto-Labeling Data for Object Detection, benchmarks auto-labeling against human annotation. We reveal how foundation models can deliver labels at near-human accuracy, whil reducing annotation costs by up to 100,000X.
geometric grey background with black gradients.

End-to-end Multimodal Platform

Improve model performance with an end-to-end multimodal annotation platform

FiftyOne is a unified data platform for multimodal and physical AI. When multimodal annotation lives alongside curation and model evaluation, your team can identify exactly where your perception or VLA model fails, surface the episodes that expose those failure modes, and route them back into the annotation pipeline, all without context-switching or round-tripping.

Multimodal annotation project management with configurable review workflows

Design and manage multi-stage multimodal annotation pipelines that give you full control over how work moves from video labeling to review and approval. Built-in project management review stages, rejection loops, and quality gates help you coordinate teams, enforce standards, and keep production datasets moving without operational friction.
FiftyOne dataset versioning interface showing previous snapshots with sample counts and a rollback option.

Standardize multimodal annotation schemas and ontologies

Give your team one source of truth for how data gets labeled. Annotation schemas set the structure, classes, and attributes for every label, and reusable ontologies let you apply the same definitions across video datasets and projects. Less ambiguity, cleaner data, and annotations that align with what your models need downstream.
FiftyOne access settings showing team members with edit, view, and tag permissions, for secure team collaboration.
FiftyOne workflow diagram: curate, annotate, generate, and evaluate multimodal data and models in a continuous loop.

Questions?
We have answers.

Get started with FiftyOne Annotation