One Camera Isn’t Enough: 9 Robotics Datasets You Can Play on One Timeline
Oct 9, 2026
•
9 min read
Author
Harpreet Sahota
Hacker-in-Residence
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.
We combine event cameras, radar, thermal, tactile pads, and lidar into nine robotics datasets, each converted into FiftyOne multimodal episodes for synchronized sensor analysis.
One of the six sequences in the DSEC sample has a mean image brightness of 26.5. The other five range from 78.5 to 92.5.
Same car, same cameras. The color images in that sequence are about a third as bright as the rest, and a pipeline tuned on the other five hasn’t seen anything like it. DSEC also records the road with a stereo pair of event cameras and ships lidar-derived disparity as ground truth, so when the color image degrades, you still have other sensors to check it against.
That is the setup for every dataset in this post. A robot that leaves the lab encounters real-world conditions such as low light, water, snow, global navigation satellite system (GNSS) dropouts, and objects it must touch. These datasets, with ground truth or robot state, help you evaluate and improve your algorithms effectively.
Each of the nine is a FiftyOne multimodal dataset: one .fo.mcap file per episode, with every sensor stream on one shared clock, making it straightforward to access and work with the data.
Dataset
Platform
Sensors beyond RGB
In this release
License
DSEC
Car
Stereo event cameras, LiDAR
6 of 53 sequences
CC BY-SA 4.0
M3ED
Car, quadrotor, and Spot
Stereo event cameras, LiDAR
3 sequences, one per robot
CC BY-SA 4.0
Edged-USLAM
Quadrotor
DAVIS346 event camera
All 13 runs
CC BY 4.0
ColoRadar
Handheld rig
Two FMCW radars, LiDAR
7 of 52 sequences
Apache 2.0
CitrusFarm
Clearpath Jackal
Thermal, red-green-NIR, depth, LiDAR
1 of 7 sequences
CC BY-SA 4.0
HapTile
UR5e arm
Vision-based tactile sensors, RGB-D
1,699 episodes
CC BY 4.0
NTNU Underwater Multi-Camera
BlueROV2 Heavy
Five monochrome cameras, barometer, rangefinder
All 8 runs
BSD 3-Clause
GrandTour
ANYmal D quadruped
HDR and depth cameras, LiDAR, GNSS/INS
3 of 49 missions
MIT
TUM RGB-D
Handheld and Pioneer robot
Kinect color and depth
47 main-table sequences
CC BY 4.0
The nine multimodal robotics datasets, with platform, sensors beyond RGB, how much of each release is included, and license.
Key takeaways
DSEC, M3ED, and Edged-USLAM record every event as a foxglove.PointCloud window of 1/30 s with x, y, z (milliseconds since the window opened), and polarity. In the spot_indoor_stairwell sequence, it runs at 63.8 Hz, a high temporal resolution.
ColoRadar, CitrusFarm, and HapTile cover radar, thermal and near-infrared (NIR), and touch: two frequency-modulated continuous-wave (FMCW) radars on a handheld rig, a thermal camera and a red-green-NIR camera on a Clearpath Jackal, and gel-pad tactile video on a UR5e gripper across 1,699 episodes.
GrandTour and NTNU Underwater Multi-Camera test assumptions about the environment. The indoor GrandTour mission ETH-2 has 0 GNSS fixes, against 44,601 in the forest mission ALB-3, and NTNU’s cameras sit behind a housing port that bends light underwater.
TUM RGB-D offers 47 sequences spanning six categories, including nine with moving people and eight designed to separate structure from texture, showcasing the datasets' broad applicability for various research needs.
Each of the nine robotics datasets is a FiftyOne dataset with media_type="multimodal" and plain sample fields such as peak_event_rate_mev_s, lighting, max_depth_m, and haptic_feedback, so you can sort and filter recordings before opening a single file.
Most of the nine robotics datasets are samples of larger releases: 6 of 53 DSEC sequences, 3 M3ED sequences, 7 of 52 ColoRadar sequences, 3 of 49 GrandTour missions, and 1 of 7 CitrusFarm sequences.
What does a camera that reports changes record?
An event camera doesn’t produce frames. Each pixel reports when its brightness changes, and nothing otherwise.
Three of the nine datasets use one, and they use it in three different places.
DSEC: the driving sequence where the color image goes dark
DSEC is a driving dataset from the University of Zurich, featuring stereo Prophesee event cameras, global-shutter color cameras, LiDAR disparity ground truth, and optical flow in some sequences, totaling 6 sequences and over 2.2 billion events.
M3ED: one sensor head on a car, a quadrotor, and a Spot
M3ED comes from the GRASP Laboratory at the University of Pennsylvania. The sensor head has a stereo pair of Prophesee EVK4 HD event cameras at 1280x720, a grayscale stereo pair, a color camera at 1280x800, an inertial unit, and an Ouster OS1-64 LiDAR. The robot changes, and the rig doesn’t. The sample contains one sequence per robot: 192 seconds and 3,631,437,863 events total. The Spot sequence alone, spot_indoor_stairwell, holds 1,883,139,312 events in 98.9 seconds.
Edged-USLAM: what an event camera records as the lights drop below 5 lux
Edged-USLAM is a DAVIS346 event camera on a quadrotor flying in a motion-capture room, recorded for an ICRA 2026 paper. Five of its 13 runs vary the lighting: 30% lighting, 60% lighting, blinking lights, strong side light, and under 5 lux. The low_lit run has the lowest mean event rate of the five: 0.53 million events per second, compared with 0.69 to 0.80 for the others. The other eight runs include fly lines, squares, circles, aggressive turns, and manual flights, with the vehicle's Vicon pose recorded alongside them.
Good to know. How are events stored in an MCAP file? Each message holds one 1/30 s window as a foxglove.PointCloud. x and y are pixel coordinates, z is the time since the window opened (ms), and polarity is 1 for a brightness increase and 0 for a decrease. The message is stamped when its window closes, so an event’s time is the stamp, minus 1/30 s, plus z. A 3D panel set to the event's frame shows each window as a volume of pixel position against time. Each camera also gets an /event-frames channel, a video render with ON events white and OFF events black on gray.
What do radar, thermal, and touch see that a color camera can’t?
ColoRadar: two radars and a lidar, carried through an underground mine
ColoRadar comes from the Autonomous Robotics and Perception Group at the University of Colorado Boulder. A handheld rig carries two FMCW radars, a TI MMWCAS-RF-EVM cascaded imaging radar and a TI AWR1843BOOST single-chip radar, next to an Ouster OS1-64 LiDAR and a Lord Microstrain 3DM-GX5-25 inertial unit. The release includes 52 sequences across seven settings: hallways, a lab, a large motion-capture space, outdoor built environments, the narrow and wide passages of an underground mine, and a fast ride along paths and roads. The sample takes the shortest sequence of each place: 7 sequences, 13m 46s, 8,263 LiDAR scans holding 382.6 million points, 4,145 cascaded radar frames, and 8,340 single-chip radar scans.
Keep it honest. The cascaded radar publishes a heatmap of 128 range by 128 azimuth by 32 elevation cells, not a point cloud. The /cascade-radar-points channel is derived from those heatmaps the way the release’s plotting tool derives them. It is not raw sensor output. The /cascade-radar-bev channel is a top-down render of the same heatmaps.
CitrusFarm: a Jackal in the orchard rows with thermal and near-infrared
CitrusFarm is a ground robot in the rows of a citrus farm, from the ARCS Lab at the University of California Riverside. A Clearpath Jackal carries a FLIR Blackfly monochrome camera, a FLIR ADK thermal camera, a Mapir Survey3 camera that records red, green, and near-infrared, a Stereolabs ZED 2i with depth, a Velodyne LiDAR, a Microstrain inertial unit, and a Piksi GPS-RTK receiver. The sample is the smallest of seven sequences: 06_14B_Jackal, 5m 03s over 357 m, with 3,024 thermal frames, 3,025 red-green-NIR frames, 2,998 LiDAR scans, and 3,024 GPS-RTK fixes. The ZED depth holds values for 47% of pixels (valid_depth_fraction = 0.4703).
HapTile: 1,699 episodes with the operator’s haptic feedback recorded
HapTile moves the question from seeing to touching. A UR5e arm with a Robotiq 2F-85 gripper works through 38 contact-rich tabletop tasks under teleoperation. Each finger carries a vision-based tactile sensor that films a marker-grid-printed gel pad, so contact shows up as markers moving. Two RGB-D cameras watch the scene, and we record the operator's haptic feedback alongside the robot state. That is 1,699 episodes, 664,045 frames, and 12.36 hours.
The feedback levels live in a haptic_feedback list on each sample. Firm appears in 1,481 episodes, soft in 575, and mild in 414. The 47 move_mobile_box episodes have no haptic channel at all.
What breaks when the robot goes underwater, indoors, or onto snow?
NTNU Underwater Multi-Camera: five cameras behind a housing port
NTNU Underwater Multi-Camera in the FiftyOne multimodal viewer
NTNU’s dataset is Ariel, a BlueROV2 Heavy from the Autonomous Robots Lab, piloted by hand through the Trondheim Fjord (6 runs) and the lab’s Marine Cybernetics Laboratory pool (2 runs). An Alphasense rig records five monochrome cameras at 720x540 and 20 Hz. The autopilot logs a second inertial unit, a barometer, a downward rangefinder, the battery, and eight thruster outputs. Each run includes a reference trajectory the authors computed with ReAqROVIO. In total, that is 63m 19s, 379,896 camera frames, and 759,943 inertial samples. The fjord runs reach 7.1 to 7.7 m deep. The two pool runs reach 0.5 and 0.6 m.
Keep it honest. The release calibrates the stereo pair underwater and the other three cameras in air only. Refraction at the housing port changes intrinsics once the vehicle is submerged, so those three carry in-air values. The depth_m track applies the release’s freshwater barometer mapping, which reads a little deep in the fjord’s salt water.
GrandTour: the same robot in a forest, a building, and snow
GrandTour is an ANYmal D quadruped from the Robotic Systems Lab at ETH Zurich carrying the Boxi sensor payload: HDR and depth cameras, a Hesai LiDAR, a tactical-grade inertial unit, and a GNSS/INS receiver, with a total station tracking a prism on the payload for reference. The release holds 49 missions. The sample carries 3 chosen to contrast: ALB-3 in a forest, ETH-2 inside the ETH main building, and SNOW-1 on the Jungfraujoch.
The indoor mission is where GNSS disappears. The GNSS field reads Working for the forest and the snow and No for ETH-2, which has 0 fixes against 44,601 and 37,201. The robot’s /legged-odometry, /lidar-odometry estimate, and /gnss-ins-odometry solution are separate pose channels, so you can overlay three estimates of the same walk and see where each one drifts from the others.
Keep it honest. The sample includes only the three HDR cameras and the upper front depth camera. The Alphasense cameras, ZED 2i, other depth cameras, Livox and Velodyne LiDARs, and the point cloud maps are in the full release but not here. For ETH-2, the release notes say its HDR timestamps are imprecise because they use the GMSL2 arrival time. That note is carried in the release_notes field.
What do you grade a trajectory against?
Every dataset in this post ships an answer key, and they are not the same kind of key. TUM RGB-D is the plainest case.
TUM RGB-D: a Kinect graded by an eight-camera motion-capture system
TUM RGB-D comes from the Computer Vision Group at the Technical University of Munich. A Microsoft Kinect records color and depth at 640x480 and 30 Hz while an eight-camera motion-capture system tracks it at 100 Hz. Three Kinects, named freiburg1, freiburg2, and freiburg3, are carried handheld and on a Pioneer robot over scenes built to test structure against texture, moving people, and object reconstruction.
This conversion includes the 47 sequences in the benchmark’s main table: 9 freiburg1, 18 freiburg2, and 20 freiburg3, totaling 48m 12s, with 81,413 color frames and 80,683 depth frames. The category field groups them into Handheld SLAM (11 sequences), 3D Object Reconstruction (11), Dynamic Objects (9), Structure vs. Texture (8), Robot SLAM (4), and Testing and Debugging (4).
Keep it honest. The validation sequences, whose ground truth the benchmark withholds, are not included. The freiburg3 sequences have no accelerometer samples because the benchmark’s archives hold none.
What each dataset grades you against
Dataset
Ground truth
DSEC
LiDAR-derived disparity for both stereo pairs, and optical flow on some sequences
M3ED
Left event camera ground-truth poses and depth
Edged-USLAM
Vicon pose of the vehicle
ColoRadar
LiDAR-inertial trajectory, plus a Vicon pose on aspen_run9
CitrusFarm
A ground-truth position track
HapTile
Robot state: end-effector pose, joints, commanded targets, and gripper
NTNU
ReAqROVIO reference trajectory
GrandTour
Total-station prism positions, plus the three odometry estimates
TUM RGB-D
Motion-capture pose of the Kinect at 100 Hz
The ground truth or reference signal that ships with each of the nine multimodal robotics datasets.
Step 1: Load a multimodal robotics dataset from Hugging Face
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
from fiftyone import ViewField as F
dsec = fouh.load_from_hub(
"Voxel51/DSEC-Sample",
name="DSEC-Sample",
persistent=True,
)
print(dsec.media_type) # multimodal
print(len(dsec)) # 6
One sample is one sequence, and filepath points to its .fo.mcap file. The datasets have no FiftyOne label fields. Every signal, including ground truth, is a channel inside the file.
Step 2: Query the sample fields before you open a file
The sample fields summarize each recording, so you can find the one you want first.
# The sequence with the darkest color images
dsec.sort_by("mean_image_brightness").first().sequence # zurich_city_09_b
# The Edged-USLAM run flown under 5 lux
edged = fouh.load_from_hub(
"Voxel51/Edged-USLAM-Event-Camera",
name="Edged-USLAM-Event-Camera",
persistent=True,
)
dark = edged.match(F("lighting") == "under 5 lux")
# HapTile episodes where the operator felt firm contact
haptile = fouh.load_from_hub("Voxel51/HapTile", name="HapTile", persistent=True)
firm = haptile.match(F("haptic_feedback").contains("firm")) # 1,481 episodes
Then launch the App and open a sample:
session = fo.launch_app(dsec)
The multimodal viewer lays out the channels as tiles with a shared playback clock, so a color frame, an event window, and a ground-truth plot move together.
Try it yourself
Each dataset above links to its Hugging Face repo and its source release. Pick the one that breaks your current pipeline, load it with the snippet in Step 1, and open a sample in the App.
Frequently asked questions
A FiftyOne multimodal dataset is a dataset with media_type="multimodal", where each sample is one recording stored as an MCAP file. The App’s viewer reads the file's channels and displays them as tiles on a single timeline. Each sample also carries ordinary fields, such as duration, num_events, or ground_truth_path_m, that you can query like any other FiftyOne field.
The nine robotics datasets have ground truth, but not as FiftyOne label fields. The ground truth, such as disparity, poses, depth, or reference trajectories, comes from each source release and lives in channels such as /ground-truth and /disparity-event. Nothing here is model output. The one derived channel is /cascade-radar-points in ColoRadar.
Four of the nine robotics datasets are complete, and five are samples. TUM RGB-D carries the benchmark’s 47 main-table sequences, NTNU carries all 8 runs, Edged-USLAM carries all 13 runs, and HapTile carries 1,699 episodes. DSEC (6 of 53 sequences), M3ED (3 sequences), ColoRadar (7 of 52), GrandTour (3 of 49), and CitrusFarm (1 of 7) are samples. The cards for those say what was left out.
The nine robotics datasets use the licenses of their source releases. NTNU is BSD 3-Clause. M3ED, DSEC, and CitrusFarm are CC BY-SA 4.0. TUM RGB-D, Edged-USLAM, and HapTile are CC BY 4.0. ColoRadar is Apache 2.0, and GrandTour is MIT.
Start with the robotics dataset that breaks your current pipeline. Low light points to Edged-USLAM, GNSS-denied walking to GrandTour ETH-2, underwater refraction to NTNU, and contact to HapTile.
Harpreet Sahota
Hacker-in-Residence
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.