Physical AI teams need data that helps models generalize, but collecting that data is often expensive, whether you’re investing in teleoperation demonstrations or real-world driving data. The challenge isn't just collecting enough data. It's also about finding the right data that improves your models.
A single multimodal recording can contain cameras, LiDAR, radar, trajectories, and other sensor streams. No single modality tells the whole story. Finding and understanding the moments that matter often means manually reviewing recordings, relying on whatever metadata happens to exist, or writing custom queries and data pipelines.
FiftyOne Search indexes your data and computes embeddings, so you can search and explore complex multimodal datasets. Find the data that matters using natural language, visual similarity, events, and metadata. And visualize the results with full temporal and spatial context.
FiftyOne Search enables teams to query and visualize multimodal data
How do you query an entire multimodal dataset?
Finding one failure tells you what happened. Improving a model often requires knowing where else it happens.
FiftyOne creates a persistent, indexed representation of your dataset so you can query across thousands of recordings rather than inspecting them one at a time. Once your data is indexed, common investigations become dataset queries instead of new data-processing projects.
Search isn't limited to what can be recognized visually. Physical behavior often emerges from signals over time: steering angle, velocity, acceleration, joint states, control inputs, or other sensor measurements.
FiftyOne lets you derive custom events from expressions over multimodal time-series signals and make those events searchable across the dataset. Once you've defined an event—such as hard braking, high steering, or a particular robot state—you can find every episode containing it and combine those results with other metadata or semantic queries.
This gives teams multiple ways into the same dataset:
Four ways to query a multimodal dataset in FiftyOne Search, and when to reach for each
Four ways to query a multimodal dataset in FiftyOne Search, and when to reach for each
Query mode
Use it when
Example
Natural language
You can describe what you are looking for
"Pedestrian crossing against a red light at night"
Visual similarity
You have an example and want more like it
Start from one failed grasp, find related grasps
Derived events
The behavior lives in signals over time
Hard braking, high steering, a specific joint state
Metadata and labels
You know the exact constraints
Model version, embodiment, object class, environment
These approaches can be combined rather than forcing every investigation through a single retrieval method.
Key takeaways
FiftyOne Search indexes multimodal datasets so you can query thousands of recordings at once instead of reviewing them one at a time.
FiftyOne Search supports four combinable query modes: natural language, visual similarity, derived events, and metadata filters.
Segment-level embeddings in FiftyOne let you retrieve the specific seconds inside a long episode that match your query, not just the episode itself.
Derived events turn multimodal time-series signals such as steering angle or joint states into searchable behaviors across the whole dataset.
Every FiftyOne Search result opens into synchronized multimodal context, so you can inspect camera, 3D, signal, and prediction streams on a shared timeline.
Find the exact moments that matter with segment-level embeddings
An episode can contain minutes or hours of data, but the behavior you care about may last only a few seconds.
FiftyOne brings segment-level embeddings to multimodal data, letting you generate embeddings over temporal segments of a recording and retrieve the specific moments within episodes that match your query. Search by natural language or visual similarity, start from an example to find semantically similar segments, or use your own embeddings models to represent domain-specific concepts.
Instead of asking whether two recordings are broadly similar, you can find the specific portions of those recordings that contain related objects, interactions, environments, or behaviors.
For physical AI, that distinction matters. A long robot demonstration might contain one failed grasp. A driving sequence might contain one unusual interaction with a cyclist. Segment-level representations let those moments become searchable without reducing the entire episode to them.
Visualize every result with multi-rate, multimodal context
Finding an interesting episode is only useful if you can understand what happened.
A camera might show that a grasp failed, while joint states and actions explain how it failed. A point cloud can reveal geometry that's ambiguous in the image. Model predictions and annotations can show whether the failure originated in perception or farther downstream.
FiftyOne connects every search result directly to synchronized multimodal visualization. Jump to a matching segment and inspect the surrounding episode across camera streams, 3D data, sensor signals, annotations, predictions, and other modalities on a shared timeline.
Configurable tiles let you customize the views that matter for the investigation, while overlays expose labels, predictions, metadata, and custom fields alongside the underlying data.
That connection also extends to analysis. Visualizations in FiftyOne are interactive, so you can move from metrics, embedding plots, and other dataset views back to the episodes and segments responsible for them.
Turn search results into better datasets
Search matters when it leads to better training and evaluation data.
FiftyOne turns search results into actionable dataset cohorts. Retrieve related samples and segments, inspect them in multimodal context, and curate the data you need for training and evaluation.
When data is expensive to acquire, this loop helps teams get more value from what they've already collected while making more informed decisions about what to collect next.
FiftyOne Search is built for physical AI data
Physical AI datasets don't fit neatly into rows of independent images. They're large, temporal, multimodal, and increasingly heterogeneous.
FiftyOne Search supports multimodal and time-series datasets, including the MCAP container format, and lets teams work with representations at the level that makes sense for the problem — from individual frames and objects to temporal segments and complete episodes.
FAQ
FiftyOne Search is a search layer for multimodal datasets that indexes your data and computes embeddings so you can find specific samples, segments, and episodes. It supports natural language, visual similarity, derived events, and metadata filters, and every result opens into synchronized multimodal visualization.
You search by content rather than by annotation. FiftyOne Search computes embeddings over your data, so a natural language description or an example sample retrieves semantically related results even when nothing in the dataset was labeled for that concept.
Segment-level embeddings are embeddings computed over temporal segments of a recording rather than over the whole recording. They let you retrieve the few seconds inside a long episode that actually match your query, which matters when a single failed grasp sits inside a multi-minute demonstration.
Yes. FiftyOne lets you define derived events from expressions over multimodal time-series signals such as steering angle, velocity, or joint states, then search for every episode containing that event and combine it with metadata or semantic queries.
FiftyOne Search supports multimodal and time-series datasets, including the MCAP container format, and lets you work at the level of frames, objects, temporal segments, or complete episodes.