One Camera Isn’t Enough: 9 Robotics Datasets You Can Play on One Timeline

Oct 9, 2026
•
9 min read
We combine event cameras, radar, thermal, tactile pads, and lidar into nine robotics datasets, each converted into FiftyOne multimodal episodes for synchronized sensor analysis.
One of the six sequences in the DSEC sample has a mean image brightness of 26.5. The other five range from 78.5 to 92.5.
Same car, same cameras. The color images in that sequence are about a third as bright as the rest, and a pipeline tuned on the other five hasn’t seen anything like it. DSEC also records the road with a stereo pair of event cameras and ships lidar-derived disparity as ground truth, so when the color image degrades, you still have other sensors to check it against.
That is the setup for every dataset in this post. A robot that leaves the lab encounters real-world conditions such as low light, water, snow, global navigation satellite system (GNSS) dropouts, and objects it must touch. These datasets, with ground truth or robot state, help you evaluate and improve your algorithms effectively.
Each of the nine is a FiftyOne multimodal dataset: one .fo.mcap file per episode, with every sensor stream on one shared clock, making it straightforward to access and work with the data.
DatasetPlatformSensors beyond RGBIn this releaseLicense
DSECCarStereo event cameras, LiDAR6 of 53 sequencesCC BY-SA 4.0
M3EDCar, quadrotor, and SpotStereo event cameras, LiDAR3 sequences, one per robotCC BY-SA 4.0
Edged-USLAMQuadrotorDAVIS346 event cameraAll 13 runsCC BY 4.0
ColoRadarHandheld rigTwo FMCW radars, LiDAR7 of 52 sequencesApache 2.0
CitrusFarmClearpath JackalThermal, red-green-NIR, depth, LiDAR1 of 7 sequencesCC BY-SA 4.0
HapTileUR5e armVision-based tactile sensors, RGB-D1,699 episodesCC BY 4.0
NTNU Underwater Multi-CameraBlueROV2 HeavyFive monochrome cameras, barometer, rangefinderAll 8 runsBSD 3-Clause
GrandTourANYmal D quadrupedHDR and depth cameras, LiDAR, GNSS/INS3 of 49 missionsMIT
TUM RGB-DHandheld and Pioneer robotKinect color and depth47 main-table sequencesCC BY 4.0
The nine multimodal robotics datasets, with platform, sensors beyond RGB, how much of each release is included, and license.

Key takeaways

  • DSEC, M3ED, and Edged-USLAM record every event as a foxglove.PointCloud window of 1/30 s with x, y, z (milliseconds since the window opened), and polarity. In the spot_indoor_stairwell sequence, it runs at 63.8 Hz, a high temporal resolution.
  • ColoRadar, CitrusFarm, and HapTile cover radar, thermal and near-infrared (NIR), and touch: two frequency-modulated continuous-wave (FMCW) radars on a handheld rig, a thermal camera and a red-green-NIR camera on a Clearpath Jackal, and gel-pad tactile video on a UR5e gripper across 1,699 episodes.
  • GrandTour and NTNU Underwater Multi-Camera test assumptions about the environment. The indoor GrandTour mission ETH-2 has 0 GNSS fixes, against 44,601 in the forest mission ALB-3, and NTNU’s cameras sit behind a housing port that bends light underwater.
  • TUM RGB-D offers 47 sequences spanning six categories, including nine with moving people and eight designed to separate structure from texture, showcasing the datasets' broad applicability for various research needs.
  • Each of the nine robotics datasets is a FiftyOne dataset with media_type="multimodal" and plain sample fields such as peak_event_rate_mev_s, lighting, max_depth_m, and haptic_feedback, so you can sort and filter recordings before opening a single file.
  • Most of the nine robotics datasets are samples of larger releases: 6 of 53 DSEC sequences, 3 M3ED sequences, 7 of 52 ColoRadar sequences, 3 of 49 GrandTour missions, and 1 of 7 CitrusFarm sequences.

What does a camera that reports changes record?

An event camera doesn’t produce frames. Each pixel reports when its brightness changes, and nothing otherwise.
Three of the nine datasets use one, and they use it in three different places.

DSEC: the driving sequence where the color image goes dark

DSEC in the FiftyOne multimodal viewer
DSEC in the FiftyOne multimodal viewer
DSEC is a driving dataset from the University of Zurich, featuring stereo Prophesee event cameras, global-shutter color cameras, LiDAR disparity ground truth, and optical flow in some sequences, totaling 6 sequences and over 2.2 billion events.

M3ED: one sensor head on a car, a quadrotor, and a Spot

M3ED in the FiftyOne multimodal viewer
M3ED in the FiftyOne multimodal viewer
M3ED comes from the GRASP Laboratory at the University of Pennsylvania. The sensor head has a stereo pair of Prophesee EVK4 HD event cameras at 1280x720, a grayscale stereo pair, a color camera at 1280x800, an inertial unit, and an Ouster OS1-64 LiDAR. The robot changes, and the rig doesn’t. The sample contains one sequence per robot: 192 seconds and 3,631,437,863 events total. The Spot sequence alone, spot_indoor_stairwell, holds 1,883,139,312 events in 98.9 seconds.

Edged-USLAM: what an event camera records as the lights drop below 5 lux

Edged-USLAM in the FiftyOne multimodal viewer
Edged-USLAM in the FiftyOne multimodal viewer
Edged-USLAM is a DAVIS346 event camera on a quadrotor flying in a motion-capture room, recorded for an ICRA 2026 paper. Five of its 13 runs vary the lighting: 30% lighting, 60% lighting, blinking lights, strong side light, and under 5 lux. The low_lit run has the lowest mean event rate of the five: 0.53 million events per second, compared with 0.69 to 0.80 for the others. The other eight runs include fly lines, squares, circles, aggressive turns, and manual flights, with the vehicle's Vicon pose recorded alongside them.
Good to know. How are events stored in an MCAP file? Each message holds one 1/30 s window as a foxglove.PointCloud. x and y are pixel coordinates, z is the time since the window opened (ms), and polarity is 1 for a brightness increase and 0 for a decrease. The message is stamped when its window closes, so an event’s time is the stamp, minus 1/30 s, plus z. A 3D panel set to the event's frame shows each window as a volume of pixel position against time. Each camera also gets an /event-frames channel, a video render with ON events white and OFF events black on gray.

What do radar, thermal, and touch see that a color camera can’t?

ColoRadar: two radars and a lidar, carried through an underground mine

ColoRadar in the FiftyOne multimodal viewer
ColoRadar in the FiftyOne multimodal viewer
ColoRadar comes from the Autonomous Robotics and Perception Group at the University of Colorado Boulder. A handheld rig carries two FMCW radars, a TI MMWCAS-RF-EVM cascaded imaging radar and a TI AWR1843BOOST single-chip radar, next to an Ouster OS1-64 LiDAR and a Lord Microstrain 3DM-GX5-25 inertial unit. The release includes 52 sequences across seven settings: hallways, a lab, a large motion-capture space, outdoor built environments, the narrow and wide passages of an underground mine, and a fast ride along paths and roads. The sample takes the shortest sequence of each place: 7 sequences, 13m 46s, 8,263 LiDAR scans holding 382.6 million points, 4,145 cascaded radar frames, and 8,340 single-chip radar scans.
Keep it honest. The cascaded radar publishes a heatmap of 128 range by 128 azimuth by 32 elevation cells, not a point cloud. The /cascade-radar-points channel is derived from those heatmaps the way the release’s plotting tool derives them. It is not raw sensor output. The /cascade-radar-bev channel is a top-down render of the same heatmaps.

CitrusFarm: a Jackal in the orchard rows with thermal and near-infrared

CitrusFarm in the FiftyOne multimodal viewer
CitrusFarm in the FiftyOne multimodal viewer
CitrusFarm is a ground robot in the rows of a citrus farm, from the ARCS Lab at the University of California Riverside. A Clearpath Jackal carries a FLIR Blackfly monochrome camera, a FLIR ADK thermal camera, a Mapir Survey3 camera that records red, green, and near-infrared, a Stereolabs ZED 2i with depth, a Velodyne LiDAR, a Microstrain inertial unit, and a Piksi GPS-RTK receiver. The sample is the smallest of seven sequences: 06_14B_Jackal, 5m 03s over 357 m, with 3,024 thermal frames, 3,025 red-green-NIR frames, 2,998 LiDAR scans, and 3,024 GPS-RTK fixes. The ZED depth holds values for 47% of pixels (valid_depth_fraction = 0.4703).

HapTile: 1,699 episodes with the operator’s haptic feedback recorded

HapTile in the FiftyOne multimodal viewer
HapTile in the FiftyOne multimodal viewer
HapTile moves the question from seeing to touching. A UR5e arm with a Robotiq 2F-85 gripper works through 38 contact-rich tabletop tasks under teleoperation. Each finger carries a vision-based tactile sensor that films a marker-grid-printed gel pad, so contact shows up as markers moving. Two RGB-D cameras watch the scene, and we record the operator's haptic feedback alongside the robot state. That is 1,699 episodes, 664,045 frames, and 12.36 hours.
The feedback levels live in a haptic_feedback list on each sample. Firm appears in 1,481 episodes, soft in 575, and mild in 414. The 47 move_mobile_box episodes have no haptic channel at all.

What breaks when the robot goes underwater, indoors, or onto snow?

NTNU Underwater Multi-Camera: five cameras behind a housing port

NTNU Underwater Multi-Camera in the FiftyOne multimodal viewer
NTNU Underwater Multi-Camera in the FiftyOne multimodal viewer
NTNU’s dataset is Ariel, a BlueROV2 Heavy from the Autonomous Robots Lab, piloted by hand through the Trondheim Fjord (6 runs) and the lab’s Marine Cybernetics Laboratory pool (2 runs). An Alphasense rig records five monochrome cameras at 720x540 and 20 Hz. The autopilot logs a second inertial unit, a barometer, a downward rangefinder, the battery, and eight thruster outputs. Each run includes a reference trajectory the authors computed with ReAqROVIO. In total, that is 63m 19s, 379,896 camera frames, and 759,943 inertial samples. The fjord runs reach 7.1 to 7.7 m deep. The two pool runs reach 0.5 and 0.6 m.
Keep it honest. The release calibrates the stereo pair underwater and the other three cameras in air only. Refraction at the housing port changes intrinsics once the vehicle is submerged, so those three carry in-air values. The depth_m track applies the release’s freshwater barometer mapping, which reads a little deep in the fjord’s salt water.

GrandTour: the same robot in a forest, a building, and snow

GrandTour in the FiftyOne multimodal viewer
GrandTour in the FiftyOne multimodal viewer
GrandTour is an ANYmal D quadruped from the Robotic Systems Lab at ETH Zurich carrying the Boxi sensor payload: HDR and depth cameras, a Hesai LiDAR, a tactical-grade inertial unit, and a GNSS/INS receiver, with a total station tracking a prism on the payload for reference. The release holds 49 missions. The sample carries 3 chosen to contrast: ALB-3 in a forest, ETH-2 inside the ETH main building, and SNOW-1 on the Jungfraujoch.
The indoor mission is where GNSS disappears. The GNSS field reads Working for the forest and the snow and No for ETH-2, which has 0 fixes against 44,601 and 37,201. The robot’s /legged-odometry, /lidar-odometry estimate, and /gnss-ins-odometry solution are separate pose channels, so you can overlay three estimates of the same walk and see where each one drifts from the others.
Keep it honest. The sample includes only the three HDR cameras and the upper front depth camera. The Alphasense cameras, ZED 2i, other depth cameras, Livox and Velodyne LiDARs, and the point cloud maps are in the full release but not here. For ETH-2, the release notes say its HDR timestamps are imprecise because they use the GMSL2 arrival time. That note is carried in the release_notes field.

What do you grade a trajectory against?

Every dataset in this post ships an answer key, and they are not the same kind of key. TUM RGB-D is the plainest case.

TUM RGB-D: a Kinect graded by an eight-camera motion-capture system

TUM RGB-D in the FiftyOne multimodal viewer
TUM RGB-D in the FiftyOne multimodal viewer
TUM RGB-D comes from the Computer Vision Group at the Technical University of Munich. A Microsoft Kinect records color and depth at 640x480 and 30 Hz while an eight-camera motion-capture system tracks it at 100 Hz. Three Kinects, named freiburg1, freiburg2, and freiburg3, are carried handheld and on a Pioneer robot over scenes built to test structure against texture, moving people, and object reconstruction.
This conversion includes the 47 sequences in the benchmark’s main table: 9 freiburg1, 18 freiburg2, and 20 freiburg3, totaling 48m 12s, with 81,413 color frames and 80,683 depth frames. The category field groups them into Handheld SLAM (11 sequences), 3D Object Reconstruction (11), Dynamic Objects (9), Structure vs. Texture (8), Robot SLAM (4), and Testing and Debugging (4).
Keep it honest. The validation sequences, whose ground truth the benchmark withholds, are not included. The freiburg3 sequences have no accelerometer samples because the benchmark’s archives hold none.

What each dataset grades you against

DatasetGround truth
DSECLiDAR-derived disparity for both stereo pairs, and optical flow on some sequences
M3EDLeft event camera ground-truth poses and depth
Edged-USLAMVicon pose of the vehicle
ColoRadarLiDAR-inertial trajectory, plus a Vicon pose on aspen_run9
CitrusFarmA ground-truth position track
HapTileRobot state: end-effector pose, joints, commanded targets, and gripper
NTNUReAqROVIO reference trajectory
GrandTourTotal-station prism positions, plus the three odometry estimates
TUM RGB-DMotion-capture pose of the Kinect at 100 Hz
The ground truth or reference signal that ships with each of the nine multimodal robotics datasets.

Step 1: Load a multimodal robotics dataset from Hugging Face

Each dataset loads from the Hugging Face Hub with one call.
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
from fiftyone import ViewField as F

dsec = fouh.load_from_hub(
    "Voxel51/DSEC-Sample",
    name="DSEC-Sample",
    persistent=True,
)

print(dsec.media_type)  # multimodal
print(len(dsec))        # 6
One sample is one sequence, and filepath points to its .fo.mcap file. The datasets have no FiftyOne label fields. Every signal, including ground truth, is a channel inside the file.

Step 2: Query the sample fields before you open a file

The sample fields summarize each recording, so you can find the one you want first.
# The sequence with the darkest color images
dsec.sort_by("mean_image_brightness").first().sequence  # zurich_city_09_b

# The Edged-USLAM run flown under 5 lux
edged = fouh.load_from_hub(
    "Voxel51/Edged-USLAM-Event-Camera",
    name="Edged-USLAM-Event-Camera",
    persistent=True,
)
dark = edged.match(F("lighting") == "under 5 lux")

# HapTile episodes where the operator felt firm contact
haptile = fouh.load_from_hub("Voxel51/HapTile", name="HapTile", persistent=True)
firm = haptile.match(F("haptic_feedback").contains("firm"))  # 1,481 episodes
Then launch the App and open a sample:
session = fo.launch_app(dsec)
The multimodal viewer lays out the channels as tiles with a shared playback clock, so a color frame, an event window, and a ground-truth plot move together.

Try it yourself

Each dataset above links to its Hugging Face repo and its source release. Pick the one that breaks your current pipeline, load it with the snippet in Step 1, and open a sample in the App.

Frequently asked questions

Headshot of Harpreet Sahota
Harpreet Sahota
Hacker-in-Residence
Harpreet is Hacker-in-Residence at Voxel51, where he turns cutting-edge AI ideas into open-source prototypes and demos that push the boundaries of deep learning. From building tools that inspire to creating content that educates, Harpreet helps the AI community level up—one wild idea at a time.
See all articles by Harpreet Sahota

Talk to an AI expert

Loading related posts...