How ThreeV Improves Training Data Quality for Utility Inspection AI with Voxel51
ThreeV builds AI that inspects power and utility infrastructure from multi-angle imagery of poles and mounted equipment. Its models depend on accurate labels for small components across millions of images. Checking those labels one CVAT job at a time meant errors surfaced only after training.
With Voxel51 as the QA layer between annotation and training, ThreeV's team now reviews labels across the whole dataset in a single view, pulls up hundreds of images of a given component at once instead of one at a time, and compares model predictions with ground truth directly on imagery in its own AWS storage. Label problems get found and fixed before the next retraining run, with far less manual review.
Category
Details
Industry
Power and utility asset inspection
Region
United States, with a globally distributed team
Team
A data science team of four and a labeling lead working in Voxel51, supported by a vendor annotation team
Data
Multi-angle inspection photos of poles and mounted equipment, with 10 to 50 images per asset and up to 150 for complex ones. Millions of images.
Models
Asset classification and defect detection
Role in the pipeline
Label QA and model evaluation layer between annotation in CVAT and model training
Stack
CVAT for annotation, MLflow, Voxel51 self-hosted in ThreeV's AWS environment, custom Python training pipelines on AWS and GCP
Key results
Thousands of labels QA'd in Voxel51 in a single view instead of one job at a time. Roughly 20 hours of review time saved per month on image loading alone. Significantly less manual QA effort.
ThreeV and Voxel51 at a glance.
Introduction: ThreeV
ThreeV builds AI that helps utilities inspect power infrastructure. Its models analyze images of poles, transmission structures, and attached equipment to identify assets and detect defects.
ThreeV's platform processes inspection imagery of transmission and distribution structures like this one. (Image: threev.ai)
ThreeV helps utility companies focus inspection and repair resources where they matter most. Instead of manually inspecting every pole, teams validate model-flagged issues and prioritize the most urgent repairs.
Behind the models is a small data science team, a labeling lead, and a vendor annotation team, working from millions of inspection images stored in the company's own AWS environment. Over the past year, that team has made Voxel51 the quality layer between annotation and training.
Challenge: Catching label errors on utility assets before they reach the model
ThreeV's models learn from annotated inspection photos of poles, conductors, insulators, and the equipment mounted on them. Each asset is photographed from several angles, and the components that matter most take up a small part of every frame. Accurate annotation is critical because distinctions such as voltage class or material carry real safety stakes. The model's output helps determine which repairs a utility puts at the front of the queue.
Previously, ThreeV exported annotations from CVAT, trained and evaluated its models, then added more data for the next iteration. Label quality was only reviewed within each completed annotation job, making it difficult to catch errors that persisted across the broader training set. A mislabeled component type or inconsistent annotation pattern could therefore go unnoticed across thousands of assets.
ThreeV wanted to:
Find missing or inconsistent labels for a given component across every asset in the dataset
Compare model predictions with ground truth on the small components where detections fail most often.
Pull up every image where a specific asset or condition appears, without zooming through dozens of photos per pole.
Work directly on imagery in its own S3 buckets, rather than copying millions of images into another tool.
For a small data science team, repeatedly loading images and assembling custom review workflows took time away from improving models.
Why ThreeV chose Voxel51
ThreeV already used CVAT for annotation. It chose Voxel51 to help its team explore datasets, QA labels, and evaluate predictions within that existing workflow.
“We chose Voxel51 because we have a limited number of data scientists here at ThreeV, and we don't want to spend most of the time loading the images and wasting our time in the basic I/O operations and QA operations.” — Jeevan, MLOps Engineer, ThreeV
The team considered building a custom viewer or using the image-logging capabilities it already had in MLflow. Voxel51 offered bulk image browsing, prediction overlays, and embeddings-based similarity search without requiring ThreeV to build those capabilities itself.
Access to existing cloud storage was another advantage. Voxel51 reads ThreeV's imagery directly from its S3 buckets, with labels and model outputs loaded alongside.
“If we had to create our own scripts or use a different application where we need to copy the data, it would definitely have led to lots of rework. With the existing data connectivity, we were directly able to visualize our model results and ground truth with filters in the Voxel51 UI.” — Jeevan, MLOps Engineer, ThreeV
Together, these capabilities gave the team a way to investigate data quality and model behavior while keeping its annotation and training tools in place.
How ThreeV uses Voxel51
ThreeV's inspection AI workflow: five steps from importing inspection imagery to labeling in CVAT, checking and evaluating in Voxel51, and retraining.
Identify labeling errors in bulk
ThreeV's labeling lead works entirely in the Voxel51 interface, without writing code. She uses filters to pull up every image carrying a given label class, and reviews them as a group rather than one job at a time.
The difference is what she can see at once. With every instance of a label side by side, an inconsistent box, a missed object, or an annotator whose work has drifted stands out in a way it never did inside a single job.
“It's very difficult to see the general picture of the quality of labels which we get. With Voxel51, it's much easier to see different patterns and to look at the data as a whole, not just specific chunks of images which were labeled.” — Polina, Labeling Lead, ThreeV
The team uses those findings to focus follow-up review on particular images or annotators' work, and to find images that should have been labeled but were not. That adds a quality check between annotation and model training, before problems in the labels become problems in the model.
Investigate difficult components and model failures
Checking the small components used to mean opening the 10 to 50 images of an asset, up to 150 for more complex assets, one-by-one and zooming in.
With Voxel51, ThreeV filters by ground-truth or model labels to bring up only the images where a given component appears. The team can confirm it has enough training examples for that component and see how they are distributed. Once a model has run, the same filter shows where its predictions match the ground truth and where they fail.
Speed matters here. In CVAT, opening a single image on the fly took around three seconds. Voxel51 loads every image that matches a filter at once, so the data science team moves through hundreds of examples in the time it used to take to open a handful.
ThreeV's transmission structure imagery in Voxel51, with ground truth and model labels for structures and wires overlaid across the sample grid. (Image: ThreeV)
The same filtered views come up in ThreeV's model review meetings. Tagging a handful of images creates a slice the team can pull up together while a model's numbers are being discussed.
“It's been helpful to bring up Voxel51 in these calls where we do huddles and look at the results of some models beyond just the metrics, to get a qualitative assessment really quickly for any arbitrary subset.” — Prajwal, CTO, ThreeV
Spend annotation budget on new examples, not duplicates
Because each asset is photographed from multiple angles, a dataset holds many near-identical shots of the same pole. Labeling all of them adds cost without improving models. ThreeV's senior ML engineer uses Voxel51's embeddings-based similarity search to find those near-duplicates before images go out for labeling, and to find the opposite: conditions that appear rarely and are under-represented in the training data.
“Voxel51's embeddings and similarity search is what we lean on to catch near-duplicates and coverage gaps before spending annotation budget.” — Anand Umashankar, Senior ML Engineer, ThreeV
This supports a more selective approach to labeling: identifying images that add useful examples, rather than repeatedly annotating similar material. ThreeV's CTO points to the same capability for the team's hardest labeling problem, conditions or equipment types that show up in a small fraction of a very large dataset, where finding more examples of a known case by similarity beats scanning for them.
Results: Thousands of labels QA'd in one view instead of one job at a time
ThreeV reports a significant reduction in the manual effort involved in reviewing labels and model results. Before Voxel51, the data science team spent most of its time loading images and running manual QA. In CVAT, each image took around three seconds to open. Voxel51 loads every image that matches a filter at once, which works out to roughly 20 hours of review time saved per month on image loading alone.
Voxel51 has also changed the team's training workflow. Instead of moving directly from annotation into training, ThreeV now examines label quality, identifies problems, and improves the data before retraining.
ThreeV's next step is to go beyond detecting individual defects and start answering more of the questions an inspector would ask about each asset. The idea is for every asset to arrive with those questions already answered by AI, along with the reasoning behind each answer, so inspectors can review and confirm the results rather than start from scratch.
Expanding the number of inspection checks a system can perform gets expensive fast, if every new check requires its own model and fine-tuning cycle. ThreeV is exploring agentic labeling, targeted fine-tuning, and similarity search as ways to broaden what the system can inspect without costs growing at the same rate. The goal is to support four to five times more inspection checks while keeping the system practical to train, review, and operate at scale.