Physical AI teams need data that helps models generalize, but collecting that data is often expensive, whether you’re investing in teleoperation demonstrations or real-world driving data. The challenge isn't just collecting enough data. It's also about finding the right data that improves your models.
A single multimodal recording can contain cameras, LiDAR, radar, trajectories, and other sensor streams. No single modality tells the whole story. Finding and understanding the moments that matter often means manually reviewing recordings, relying on whatever metadata happens to exist, or writing custom queries and data pipelines.
Voxel51 Search indexes your data and computes embeddings, so you can search and explore complex multimodal datasets. Find the data that matters using natural language, visual similarity, events, and metadata. And visualize the results with full temporal and spatial context.
Voxel51 Search enables teams to query and visualize multimodal data
How do you query an entire multimodal dataset?
Finding one failure tells you what happened. Improving a model often requires knowing where else it happens.
Voxel51 creates a persistent, indexed representation of your dataset so you can query across thousands of recordings rather than inspecting them one at a time. Once your data is indexed, common investigations become dataset queries instead of new data-processing projects.
Search isn't limited to what can be recognized visually. Physical behavior often emerges from signals over time: steering angle, velocity, acceleration, joint states, control inputs, or other sensor measurements.
Voxel51 lets you derive custom events from expressions over multimodal time-series signals and make those events searchable across the dataset. Once you've defined an event—such as hard braking, high steering, or a particular robot state—you can find every episode containing it and combine those results with other metadata or semantic queries.
This gives teams multiple ways into the same dataset:
Four ways to query a multimodal dataset in Voxel51 Search, and when to reach for each
Four ways to query a multimodal dataset in Voxel51 Search, and when to reach for each
Query mode
Use it when
Example
Natural language
You can describe what you are looking for
"Pedestrian crossing against a red light at night"
Visual similarity
You have an example and want more like it
Start from one failed grasp, find related grasps
Derived events
The behavior lives in signals over time
Hard braking, high steering, a specific joint state
Metadata and labels
You know the exact constraints
Model version, embodiment, object class, environment
These approaches can be combined rather than forcing every investigation through a single retrieval method.
Key takeaways
Voxel51 Search indexes multimodal datasets so you can query thousands of recordings at once instead of reviewing them one at a time.
Voxel51 Search supports four combinable query modes: natural language, visual similarity, derived events, and metadata filters.
Segment-level embeddings in Voxel51 let you retrieve the specific seconds inside a long episode that match your query, not just the episode itself.
Derived events turn multimodal time-series signals such as steering angle or joint states into searchable behaviors across the whole dataset.
Every Voxel51 Search result opens into synchronized multimodal context, so you can inspect camera, 3D, signal, and prediction streams on a shared timeline.
Find the exact moments that matter with segment-level embeddings
An episode can contain minutes or hours of data, but the behavior you care about may last only a few seconds.
Voxel51 brings segment-level embeddings to multimodal data, letting you generate embeddings over temporal segments of a recording and retrieve the specific moments within episodes that match your query. Search by natural language or visual similarity, start from an example to find semantically similar segments, or use your own embeddings models to represent domain-specific concepts.
Instead of asking whether two recordings are broadly similar, you can find the specific portions of those recordings that contain related objects, interactions, environments, or behaviors.
For physical AI, that distinction matters. A long robot demonstration might contain one failed grasp. A driving sequence might contain one unusual interaction with a cyclist. Segment-level representations let those moments become searchable without reducing the entire episode to them.
Visualize every result with multi-rate, multimodal context
Finding an interesting episode is only useful if you can understand what happened.
A camera might show that a grasp failed, while joint states and actions explain how it failed. A point cloud can reveal geometry that's ambiguous in the image. Model predictions and annotations can show whether the failure originated in perception or farther downstream.
Voxel51 connects every search result directly to synchronized multimodal visualization. Jump to a matching segment and inspect the surrounding episode across camera streams, 3D data, sensor signals, annotations, predictions, and other modalities on a shared timeline.
Configurable tiles let you customize the views that matter for the investigation, while overlays expose labels, predictions, metadata, and custom fields alongside the underlying data.
That connection also extends to analysis. Visualizations in Voxel51 are interactive, so you can move from metrics, embedding plots, and other dataset views back to the episodes and segments responsible for them.
Turn search results into better datasets
Search matters when it leads to better training and evaluation data.
Voxel51 turns search results into actionable dataset cohorts. Retrieve related samples and segments, inspect them in multimodal context, and curate the data you need for training and evaluation.
When data is expensive to acquire, this loop helps teams get more value from what they've already collected while making more informed decisions about what to collect next.
Voxel51 Search is built for physical AI data
Physical AI datasets don't fit neatly into rows of independent images. They're large, temporal, multimodal, and increasingly heterogeneous.
Voxel51 Search supports multimodal and time-series datasets, including the MCAP container format, and lets teams work with representations at the level that makes sense for the problem — from individual frames and objects to temporal segments and complete episodes.
FAQ
Voxel51 Search is a search layer for multimodal datasets that indexes your data and computes embeddings so you can find specific samples, segments, and episodes. It supports natural language, visual similarity, derived events, and metadata filters, and every result opens into synchronized multimodal visualization.
You search by content rather than by annotation. Voxel51 Search computes embeddings over your data, so a natural language description or an example sample retrieves semantically related results even when nothing in the dataset was labeled for that concept.
Segment-level embeddings are embeddings computed over temporal segments of a recording rather than over the whole recording. They let you retrieve the few seconds inside a long episode that actually match your query, which matters when a single failed grasp sits inside a multi-minute demonstration.
Yes. Voxel51 lets you define derived events from expressions over multimodal time-series signals such as steering angle, velocity, or joint states, then search for every episode containing that event and combine it with metadata or semantic queries.
Voxel51 Search supports multimodal and time-series datasets, including the MCAP container format, and lets you work at the level of frames, objects, temporal segments, or complete episodes.