FiftyOne's Video Timeline Catches the One Track a Human Annotator Dropped
Oct 7, 2026
•
5 min read
Author
Adonai Vera
Adonai Vera is a Machine Learning Engineer & DevRel at Voxel51 with over 7 years of experience building computer vision and machine learning models using TensorFlow, Docker, and OpenCV. Adonai started as a software developer, moved into AI, led teams, and served as CTO. Today, he connect code and community to build open, production-ready AI — making technology simple, accessible, and reliable. LinkedIn | GitHub
FiftyOne 1.22.0's new video Explore timeline lets you scrub a video sample's per-object tracks without ever opening Annotate. On a real dashcam clip it surfaces one vehicle, out of 195 tracked objects across all ten quickstart-video clips, that a human annotator lost for 54 straight frames and picked back up under the same id once traffic started moving again.
A new drawer under the video, not a new panel
Opening a video sample in FiftyOne's Explore mode used to give you a player: play, pause, step. FiftyOne 1.22.0 adds a real timeline docked under that player, the same shared-clock timeline the multimodal viewer already uses for MCAP episodes, now wired up for plain video datasets too. Every track in the sample's frame labels gets its own row, drawn as one interval bar per contiguous span the object was actually detected in. You never have to switch to Annotate mode to see it.
That distinction matters for a very ordinary reason: Explore is read only. A data engineer auditing annotation quality, or a reviewer checking a vendor's delivery, can scrub straight through a clip's object tracks without any risk of nudging a box and creating an edit nobody asked for.
Key takeaways
FiftyOne 1.22.0 adds a video timeline to Explore mode that draws one interval bar for each continuous span a tracked object appears in, so gaps in a track are visible without opening Annotate mode.
Because Explore mode is read-only, reviewers can audit video annotation quality in the FiftyOne video timeline without accidentally editing labels.
In the FiftyOne quickstart-video dataset, only 2 of 195 human-annotated tracks (1.0%) contain a gap, and the overall dropout rate is 0.52%.
The largest track gap in quickstart-video is a vehicle missing from the ground truth for 54 consecutive frames (1.80 seconds) while still on screen. It then picks back up under the same track index.
A short script that groups frame-level detections by label and track index can find track gaps in any FiftyOne video dataset with tracked detections.
Let's test it on the quickstart-video dataset
We loaded the FiftyOne Zoo's quickstart-video dataset (ten real dashcam-style clips, dense human-drawn detections, three classes: vehicle, road sign, person) and asked a simple question the timeline is built to answer at a glance: does every tracked object stay tracked for its whole time on screen, or does the ground truth itself have gaps?
All ten quickstart-video clips, real dashcam footage with dense per-frame detections.
What the FiftyOne video timeline actually shows
Opening the clip that turned out to matter (a slow, nearly stationary moment at a downtown Los Angeles intersection) puts you in the real video Explore surface: the player, a transport bar with play, step, volume and playback speed, and a seconds ruler.
The real video Explore surface: player, transport controls, and the new seconds ruler.
Below that ruler, FiftyOne now renders one interval bar per contiguous span each tracked object was actually drawn in, not just its first-to-last frame. We confirmed this directly: the vehicle at the center of this post renders as two separate bars in the same lane, one from 2.67s to 2.77s and a second from 4.57s to 5.34s, the exact shape of a track that drops out mid-clip and comes back under the same id.
The measurement: how many tracks have gaps?
We wrote a small script that reads every sample's frame-level detections, groups them by (label, track index), and for each track counts frames strictly between its first and last appearance where that same track index is simply missing. No model, no heuristic: this is exactly what a human annotator drew, read back literally.
Track gaps by class across all 10 quickstart-video clips: only 2 of 195 human-annotated tracks have any gap.
Label
Tracks
Tracks with a gap
Frames spanned
Gap rate
Vehicle
109
1
7565
0.71%
Road sign
62
1
2731
0.18%
Person
24
0
1108
0.00%
Track gaps by class across all 10 quickstart-video clips: only 2 of 195 human-annotated tracks have any gap.
Across all 195 tracks in all 10 clips, only 2 have any gap at all: 1.0%. The overall dropout rate, gap frames over every frame every track should have covered, is 0.52%. This is, in other words, a very clean dataset. That is not a disappointing result, it is what makes the one real exception worth looking at.
That exception: a vehicle track (frames 81 to 160, an 80-frame span) is missing for 54 consecutive frames, a 67.5% dropout rate for that one track, 1.80 seconds at this clip's 29.97 fps.
Frame 82: tracked. Frame 110: the same car, still visible to a human eye, missing from the ground truth. Frame 138: tracked again, same track id.
The three frames above are not staged. They are pulled straight from the clip with OpenCV, and the orange box is the dataset's own ground-truth bounding_box for that exact frame, drawn by our script, not by hand. The car barely moves on screen across these 56 frames, the ego vehicle is nearly stopped at a green light, and the box's own width and height are identical before and after the gap (0.025 x 0.044, normalized). This was not an object that left the frame or changed shape enough to confuse a tracker. An annotator's box simply stopped being drawn for 54 frames and picked back up on the same object.
Build a track gap check for your own video dataset
The recipe generalizes to any FiftyOne video dataset with a tracked detections field (any field where labels carry an index):
tracks = {}
for frame_number, frame in sample.frames.items():
for det in frame[label_field].detections:
if det.index is None:
continue
tracks.setdefault((det.label, det.index), []).append(frame_number)
# for each track: gap_frames = frames in [first, last] not present
# sort by longest contiguous gap, jump straight to the worst offender
# in the App's video Explore timeline, no Annotate mode required
The same shape works whether the field is on a dashcam dataset like this one, a sports-analytics feed, or a wildlife camera trap: the gap check does not care what the boxes are of, only whether the annotator's own index stayed continuous.
Where this matters outside a demo dataset
Autonomous driving annotation QA: vendors delivering per-frame boxes for dashcam or AV log footage are paid for continuous tracks. A 1.8 second dropout in the middle of a clip, on an object that never left the frame, is exactly the kind of defect a spot check misses and a full frame-by-frame review is too slow to catch at scale.
Traffic and retail camera analytics: dwell-time and counting pipelines that key off a persistent track id will silently double-count or drop an object across a gap like this one, since a system that treats index continuity strictly sees the reappearance as a new object arriving.
Any outsourced video annotation program: sports analytics, security footage, camera-trap wildlife data, anywhere boxes are drawn by a labeling vendor rather than a model, this same script is the fast way to find the handful of tracks worth escalating before the delivery ships to training.
Open the video sample in Explore mode in FiftyOne 1.22.0 or later. The timeline under the player shows one row per tracked object, with a bar for each span the object is labeled in. Explore mode is read-only, so you can't accidentally edit a box.
Group frame-level detections by label and track index, then count the frames between each track's first and last appearance where that index is missing. Sort by the longest gap and open the worst track in the FiftyOne video timeline.
A track gap is a run of frames between an object's first and last labeled appearance where its track index is missing. In quickstart-video, one vehicle is still visible but has no ground-truth box for 54 frames, then resumes under the same ID.
Track gaps break any pipeline that relies on a persistent track ID. Dwell-time and counting systems can double-count or drop the object. For vendor-labeled video, a gap is the kind of defect spot checks miss.
Adonai Vera
Adonai Vera is a Machine Learning Engineer & DevRel at Voxel51 with over 7 years of experience building computer vision and machine learning models using TensorFlow, Docker, and OpenCV. Adonai started as a software developer, moved into AI, led teams, and served as CTO. Today, he connect code and community to build open, production-ready AI — making technology simple, accessible, and reliable.